Fetching from the wire…
Public story · 2026-02-22 · source-backed
Each link below shares sources, entities, or timing with this story.
ARC-AGI-2 Reasoning Race. Gemini 3.1 Pro hit 77.1% and Claude Opus 4.6 hit 68.8% — both more than doubling their predecessors' scores in a single generation. The entire field moved ~10 points over the prior two years; each model jumped 30-40 points in one release. The top thre...
42. Vellum AI — Opus 4.6 Benchmarks 43. Arcee AI — Trinity Large 44. arXiv — PCAS 45. arXiv — SpargeAttention2 46. arXiv — CUWM 47. arXiv — VESPO 48. arXiv — SAGE 49. The Register — TDD for AI Agents 50. Latent Space — Anita TDD 51. builder.io — TDD + AI
This is the cleanest experimental result I've seen in weeks, and it explains a class of debugging pain I've hit personally. Researchers held weights, test cases, decoding parameters and seeds fixed on BFCL v4 and changed exactly one thing: the serving adapter. The tool-call sc...
While the capability stories pile up, here's the counterweight. As SWE-bench Verified scores cluster near saturation on July leaderboards, an enhanced analysis (SWE-Bench+, on the AIware 2026 benchmark track) found 60.83% of commonly resolved issues contain solution leakage ri...
Largest open-source sparse MoE model: 400B parameters, 13B active per token, trained on 17 trillion tokens. Novel SMEBU for load balancing. Zero loss spikes with Muon optimizer across 17T tokens — a remarkable training stability achievement. arXiv | Arcee Blog
First interactive reasoning benchmark. Top agent (StochasticGoose) scored 12.58% vs. humans. "Intelligence is efficiency." Agents struggle to convert environmental feedback into coherent strategies. Full launch March 25. ARC Prize
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.