Fetching from the wire…
Public story · 2026-03-05 · source-backed
ZeroDayBench (2603.02297) — GPT-5.2, Claude Sonnet 4.5, and Grok 4.1 all fail at autonomous zero-day vulnerability discovery. Reality check: the CyberStrikeAI threat is automation of known exploits, not novel vulnerability discovery. ICLR 2026 Workshop.
tau-Knowledge (2603.04370) — Frontier models achieve only 25.5% pass rate on fintech customer support with ~700 interconnected knowledge documents. Unstructured knowledge retrieval + policy compliance remains a major unsolved challenge.
SkillCraft (2603.00718) — Tool composition and reuse reduces token costs by 80%. Skill caching is the next efficiency frontier for agent frameworks.
Each link below shares sources, entities, or timing with this story.
Netlify published an AXIS-framework evaluation on August 14 that I've been thinking about all day. Same task, 11 models, three runs each, scored on functional correctness rather than aesthetics. The task was deliberately boring: a static one-page coffee shop site with hours, a...
Four frontier models. Five sealed engineering problems. The result everybody will quote is that Claude Fable 5 won. The result that should actually change how you work is buried three-quarters down the page. JuliaHub published an evaluation on July 30 running four frontier mod...
An automated framework evaluated GPT, Gemini, Claude and Grok on 85 algorithmic C# tasks derived from HumanEval, producing 340 solutions scored on three independent axes: functional correctness via unit tests, static quality via Roslyn AST analysis, and runtime efficiency via...
Four configurable reasoning levels (low, medium, high, xhigh), positioned for long-running agents and visual work. (xAI) The pricing undercuts frontier Claude and GPT tiers on Bedrock, which matters if you route agent traffic by cost. Watch the context pricing structure, thoug...
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill. SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-...
OpenAI released GPT-5.4 in Standard, Thinking, and Pro variants. Headline capabilities: native computer-use (75.0% on OSWorld-Verified, surpassing human 72.4%), 1M token context, and first-ever "compaction" support for longer agent trajectories. The Tool Search API is the buil...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.