Fetching from the wire…
Agents2026-08-19 · source-backed
StartupBench (arXiv 2608.17800) inverts benchmark construction. Instead of researcher-invented tasks, the authors studied AI startup products with demonstrated market adoption, their workflows, and their users, then translated those into complete deliverable-oriented tasks with fine-grained rubrics. Under a unified harness, the strongest model completes roughly 30% end to end, while making substantial partial progress on many tasks. The named failure sources: complex instruction following and domain-specific expertise. Neither is a raw capability problem. Both are things you fix with rubric-shaped decomposition and injected domain context, which means the missing 70% is addressable by product engineering rather than by waiting for the next model.
Each link below shares sources, entities, or timing with this story.
Under one identical GUI-MCP harness on OSWorld-MCP's 309 tasks, the same MCP tools improved a reasoning model by +4.0pp and degraded a non-reasoning model by -5.9pp (arXiv 2608.03327). The authors call it the "adoption gap": the reasoning model used a tool on just 55 of 309 ta...
Learned KV eviction has a soft-to-hard mismatch: training uses differentiable gates that attenuate contributions, but inference only saves memory when entries are physically removed (arXiv 2608.23296). A controlled 2x2x2 over attention type, learned gating and positional encod...
arXiv 2608.06370 evaluated models emitting code that calls tools against JSON-schema tool calling on BFCL v4. PTC matched or exceeded the baseline in 11 of 14 models, with the GPT-5.6 family up 10.6%, and held stable under parallel execution in 13 of 14. Under context degradat...
arXiv 2608.03609 formalizes agentic systems over relational data as Stateful Tool-Enabled Agentic Deployments and proves verification against First-Order CTL specs is undecidable. Under a finite-domain restriction it becomes PSPACE-complete, but only if renaming opaque identif...
Four frontier models. Five sealed engineering problems. The result everybody will quote is that Claude Fable 5 won. The result that should actually change how you work is buried three-quarters down the page. JuliaHub published an evaluation on July 30 running four frontier mod...
Most benchmarks test single-episode solving, and memory benchmarks test fact retention. Neither checks procedural reuse, whether an agent can convert a solved session into a reusable search/debug/verify routine (arXiv). Under a Train/Extract/Test protocol with held-out tasks,...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.