Fetching from the wire…
Research2026-08-20 · source-backed
40+ experts and 31,000+ human hours built 60 project-level research tasks across 11 domains, designed so guidance is progressively withdrawn within the same project. Across 18 agent-model configurations: 50.91 with full methodological guidance, 29.10 when only the method is specified, 26.62 when the agent picks the method itself. (arXiv 2608.17271) Half the score was the human's methodology. Same conclusion as the Princeton study from a completely different direction.
Each link below shares sources, entities, or timing with this story.
Writer launched Palmyra X6 on August 13 with a number that should reset how you think about agent COGS: 52% lower average cost, 48% better speed, 10% better quality. The model is a post-training variation of Z.ai's open-source GLM-5.2. A US enterprise SaaS vendor built its fla...
Cline released @cline/sdk on May 13, an open-source TypeScript agent runtime that powers their CLI, VS Code, and JetBrains extensions. Running claude-opus-4.7, Cline CLI scores 74.2% on Terminal-Bench 2.0. Claude Code on the same model: 69.4%. Same model. Different harness. Al...
Stanford's Denisov-Blanch group built a maturity model for AI adoption scored entirely from artifacts already in version control, applied it to 441 repositories, and found something I've been assuming without evidence. RAMP is a four-level model derived only from committed AI...
Built with 250-plus industry experts, ALE evaluates agents on 960 expert-authored, deterministically-scored workflows across 55 industries. Agents average 26% across all tiers but 2.6% on the "last-exam" tier. The kicker: Codex with GPT-5.5 scores 82% on Terminal-Bench and 0%...
Somebody finally measured the thing everyone complains about, and the numbers are worse than the vibes. A Level1Techs writeup that hit 384 points and 144 comments on Hacker News captured full-vocabulary logits and computed KL divergence in FP64 to trace exactly where local inf...
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.