Fetching from the wire…
Public story · 2026-02-23 · source-backed
ARC-AGI-2 Reasoning Race. Gemini 3.1 Pro hit 77.1% and Claude Opus 4.6 hit 68.8% — both more than doubling their predecessors' scores in a single generation. The entire field moved ~10 points over the prior two years; each model jumped 30-40 points in one release. The top three coding models (Opus 4.6 80.8%, Codex 80.0%, Gemini 3.1 Pro 80.6%) are in a statistical tie on SWE-bench. Source: Vellum AI
Arcee Trinity Large: Open 400B Sparse MoE. Largest open-weight US model. 400B total, 13B active per inference. MMLU 87.2 (vs Llama 4 Maverick 85.5). Trained in 33 days for ~$20M. Apache 2.0. 2-3x faster inference due to extreme sparsity. Source: Arcee AI Blog | arXiv
PCAS: Policy Compiler for Agent Security. Deterministic enforcement via Datalog-derived policy language. Improved compliance from 48% to 93% with zero violations. Fills the "enforcement gap" that threat models identified but didn't solve. Source: arXiv
SpargeAttention2: 95% Sparsity, 16.2x Speedup. Hybrid masking for robust sparse attention. Applicable for video generation and long-context inference. Source: arXiv
CUWM: Computer-Using World Model. Microsoft's desktop world model predicts UI state before acting. Test-time action search for GUI agents. Source: arXiv
VESPO: Stable Off-Policy LLM Training. Top paper on HuggingFace (152 upvotes). Stable at 64x staleness. Applicable for RLHF/RLAIF. Code on GitHub. Source: arXiv
SAGE: Self-Aware Guided Efficient Reasoning. Reasoning models know when to stop thinking, but sampling obscures it. Practical technique for reducing inference costs. Source: arXiv
Each link below shares sources, entities, or timing with this story.
42. Vellum AI — Opus 4.6 Benchmarks 43. Arcee AI — Trinity Large 44. arXiv — PCAS 45. arXiv — SpargeAttention2 46. arXiv — CUWM 47. arXiv — VESPO 48. arXiv — SAGE 49. The Register — TDD for AI Agents 50. Latent Space — Anita TDD 51. builder.io — TDD + AI
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
This one annoyed me, because I've been running the losing pattern. SWE-QA (arXiv 2608.01507) compares the sub-agent grep pattern that Claude Code, Codex and Antigravity all ship by default against a pre-built semantic index over the same repository. Semantic search answered 65...
36. arXiv — CUWM 37. arXiv — IntentCUA 38. arXiv — SpargeAttention2 39. arXiv — Arcee Trinity Large 40. ARC Prize — ARC-AGI-2 41. Arcee Blog ---
The ARC Prize Foundation dropped ARC-AGI-3 on March 25 and the results broke my mental model of how AI capability scales. Symbolica's Arcgentica framework scored 36.08% (113 of 182 playable levels, 7 of 25 games completed) using Claude Opus 4.6 as its backbone. Cost: $1,005. F...
Three moves, two days, no coordination between them. August 10–11: GitHub shipped Ollama as a BYOK provider inside Copilot for JetBrains (GitHub Changelog). Unsloth released Unsloth Desktop with a command literally named unsloth start claude, which points Claude Code and Codex...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.