Fetching from the wire…
Research2026-08-20 · source-backed
A Princeton-led study with Stanford, Berkeley, Johns Hopkins, Toronto, Georgetown and the UK AI Security Institute handed Claude Opus 4.8 and GPT-5.6 Sol Ultra the central research question from unpublished NeurIPS 2026 submissions. The original authors graded the output as reviewers and rejected both, citing weak experimental design, unsupported conclusions and impenetrable prose. Neither agent spent its full budget. (MIT Technology Review) Kirgis and Kapoor argue the failure was judgment, not engineering: the agents could run experiments and write LaTeX, they just couldn't decide which hypotheses deserved compute.
Each link below shares sources, entities, or timing with this story.
MIT Technology Review covers a Stanford/MIT effort (Anka Reuel, Shayne Longpre) analyzing 24,521 donated conversations across 52 models from 2023-2025 against vendors' own usage reports. Because Anthropic filters for work-related use, nearly half of real conversations would be...
If you have a CLAUDE.md, you're in scope. Today. arXiv 2607.14611 (cs.CR, filed July 16) evaluates prompt injection planted in the persistent memory files that agentic coding systems write and re-read across sessions. The researchers tested both Anthropic's Claude Code and Ope...
GPT-5.6 Sol Ultra tops out at 91.9%. The public leaderboard is led by Codex CLI plus GPT-5.5 at 83.4%, with Claude Code plus Opus 4.8 the top usable Claude pairing at 78.9%, and Gemini CLI plus Gemini 3.1 Pro at 70.7% (Morph). There are now roughly 35 actively maintained CLI c...
Three independent companies converged on the same architectural insight within days. That's not a coincidence. That's a pattern. Cursor 3 launched April 2 with a complete IDE rebuild centered on an Agents Window for parallel AI fleets. The /best-of-n command runs the same task...
Microsoft announced Critique on March 30. Here's how it works: when you use M365 Copilot Researcher, GPT drafts the initial research response. Then Claude reviews it for accuracy, completeness, and citation quality. You only see the final result after both models have had thei...
Netlify published an AXIS-framework evaluation on August 14 that I've been thinking about all day. Same task, 11 models, three runs each, scored on functional correctness rather than aesthetics. The task was deliberately boring: a static one-page coffee shop site with hours, a...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.