Fetching from the wire…
Research2026-08-27 · source-backed
FrontierChallenge released 97 of 300 end-to-end scientific workflows across quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science and electrochemistry, each specifying a bundle of required deliverables rather than a final answer (arXiv 2608.24979). Twelve frontier models across three scaffolds topped out at 20 of 97 tasks. Partial credit is actively misleading: analytical chemistry and electrochemistry scored 87.6 and 94.9 average with pass rates of 4% and 0%. The 75.5% false-completion rate is the number to hold onto, because it means an agent's own report of success carries close to no information at these difficulty levels.
Each link below shares sources, entities, or timing with this story.
This one annoyed me, because I've been running the losing pattern. SWE-QA (arXiv 2608.01507) compares the sub-agent grep pattern that Claude Code, Codex and Antigravity all ship by default against a pre-built semantic index over the same repository. Semantic search answered 65...
SkillSentry (arXiv 2608.09253) targets the gap where an agent completes a task under skill guidance then fails the same task on a repeat run. It defines a DSL for runtime guidance, initializes it from skill specs plus insights mined from historical successful *and failed* trac...
Four propositions, and the second is the one I'd print out. Instruction, permission enforcement, sandboxing and OS isolation are four distinct layers, only two are enforced, and conflating them is the most common cause of losing control. The others: capability without a define...
Thinkingbox is an MCP-compatible sandbox with isolated sessions, full execution traces, and outcome evaluation against terminal backend state, carrying 507 policy-conditioned workflows across retail, hospitality, auto insurance, neobank internal IT and consulting support (arXi...
Deng et al. built 120 real-case-grounded tasks across 20 business scenes in six financial domains, running four self-evolving scaffolds on a shared Qwen3.7-Max backbone against paired non-evolving controls. Letta posted the highest evolved score (91.65) and fewest compliance i...
Microsoft's July 23 release targets a genuine gap: harness-based agents like Claude Code and Codex drive multi-turn reasoning, tool use, and external system access but were hard to train end-to-end with standard open RL infrastructure. The trick is decoupling training from inf...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.