Fetching from the wire…
Public story · 2026-08-29 · source-backed
Alibaba International's Accio team open-sourced 107 tasks (53 CLI, 28 browser, 16 file, 10 API/MCP) running against fourteen offline replicas of real business software in a fresh container per task, with verifiers inspecting mock-service state rather than the transcript. Claude Opus 5 leads at 65/107 (60.7%) using 52.5 steps, 7.8 minutes and 2.05M tokens per task, ahead of Opus 4.8 at 56/107, with Qwen 3.8 Max and DeepSeek V4 Pro tied at 53/107. State-inspecting verifiers are the design decision every agent benchmark should copy. (GitHub)
Each link below shares sources, entities, or timing with this story.
Accio open-sourced it August 2, now 1,042 stars: 107 tasks across 14 containerized mock services replicating Alibaba, Shopify, and FreightOS (GitHub). Task mix is 53 CLI, 28 browser, 16 file ops, 10 API/MCP, split 65 text-only / 20 browser-text / 22 vision-requiring. Twelve mo...
A GitHub repo cataloging Claude Code tips doesn't normally warrant a top story. But shanraisshan/claude-code-best-practice at 53.4K stars isn't a tips list anymore. It's the de facto reference for how an entire generation of developers is learning to work with AI coding agents...
This is the week's most useful number, and it took an actual experiment to produce it rather than a launch post. Together AI ran 113 DeepSWE tasks at 4 trials per config: 452 GLM-5.3 rollouts and 448 GLM-5.3-Flash rollouts. GLM-5.3 scored 69.0% pass@1 at $3.99 per rollout. Fla...
affaan-m/ECC (36.3k forks, MIT) bundles 67 agents, 284 skills, 94 legacy command shims, and "instincts", patterns learned from prior sessions with confidence scores that auto-recall when relevant, plus a .ecc/memory/ markdown vault that's explicitly cross-harness, so context s...
The leaderboard says first place. The methodology says you should check your own bill. Qwen3.8 Max now ranks first on Artificial Analysis' agentic index, scoring 86.1 on OSWorld-Verified ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0, priced at $2.00/M input and $6.00/M...
DeepClaude hit 470 points on Hacker News. It swaps Claude Code's API backend to DeepSeek V4 Pro while preserving the full agent loop: file editing, bash execution, git tooling, the whole workflow. DeepSeek V4 Pro scores 96.4% on LiveCodeBench at a fraction of Anthropic's prici...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.