Fetching from the wire…
Public story · 2026-02-12 · source-backed
Each link below shares sources, entities, or timing with this story.
20. Schneier — Anthropic and the Pentagon 21. Interconnects — OLMo Hybrid Analysis 22. ARC Prize — ARC-AGI-3 23. EpicAI.pro — Kent C. Dodds
15. CNBC — Altman Pentagon Admission 16. simonwillison.net — Cognitive Debt 17. Lenny's Newsletter — Rauch v0 18. ARC Prize — ARC-AGI-3 19. Semafor — Sharma Resignation
GPT-5.4 scores 0.26%. Opus 4.6 scores 0.25%. Grok-4.20 scores 0.00%. Humans score 100%. The Decoder covered the ARC-AGI-3 launch on March 25, and the results make every "AGI is here" claim look premature. François Chollet launched ARC-AGI-3 at Y Combinator HQ alongside a fires...
43.3% on Frontier-Bench v0.1. Opus 4.8 scored 18.7%. That's not an incremental bump, that's the same benchmark with a different shape of answer. Anthropic released Claude Opus 5 on July 24 at $5/$25 per million input/output tokens, exactly half of Fable 5's $10/$50, while matc...
Boris Cherny told Simon Willison the most exciting thing about Opus 5 isn't the evals, it's the injection resistance, and that it's "a bit buried in the system card" on page 73. The numbers: attacker success within 15 attempts on the Gray Swan indirect injection benchmark fell...
No new posts today. Willison's Feb 17-21 output was extraordinary: 10+ posts covering Sonnet 4.6, GGML/HuggingFace merger, SWE-bench analysis (Opus 4.5 leads at 76.8%, Chinese models dominate top 10), and the Karpathy "Claws" amplification. His Showboat ecosystem — Rodney, Cha...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.