Fetching from the wire…
Public story · 2026-08-07 · source-backed
The leaderboard says first place. The methodology says you should check your own bill.
Qwen3.8 Max now ranks first on Artificial Analysis' agentic index, scoring 86.1 on OSWorld-Verified ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0, priced at $2.00/M input and $6.00/M output. Its Intelligence Index score of 56 puts it level with Claude Opus 4.8 (max) and ahead of every model Google, Meta and xAI ship. 507 points and 320 comments on HN.
Buried in the data: it averages 64 turns on GDPval-AA. Qwen3.7 Max averaged 14. That's a ~4.5x turn-count tax, and a per-token price comparison shows you exactly none of it. If your agent loop reloads context each turn (most do), turn count is closer to a multiplier on your actual spend than a footnote.
Three other findings today say the same thing from different angles. RealReplicaBench (1,038 stars in 5 days) ran 12 models across 107 long-horizon tasks in high-fidelity stateful clones of eight commerce and logistics platforms. Claude Opus 5 led at 66/107 (61.7%) on the Accio harness and 60/107 (56.1%) on OpenClaw. Same model, same tasks, 5+ points of swing from swapping the scaffold. DCAS found that open coding models fine-tuned on OpenHands trajectories degrade substantially under any other CLI scaffold, while untrained base models show no such divergence, pinning the load-bearing variable on planning structure. And the single largest sentiment signal across tracked subreddits today was a meme mocking benchmark charts at 5,938 upvotes with only 56 comments. Low comment ratio means consensus, not argument. Nobody's defending vendor evals.
The arithmetic to run before you swap models: measure average turns per completed task on your workload, multiply by your average context size, multiply by input price. Then compare. I'd bet on the cheaper-per-token model losing that comparison more often than the leaderboards imply, and I'd also bet most teams have never run it.
Related and worth pricing in: DeepSeek emailed API users on August 6 warning of a "significant" price increase, citing demand beyond platform capacity and unsustainable server costs. V4-Flash currently sits at $0.14/M input and $0.28/M output, the floor anchoring the whole cheap-inference market. Second pricing change in under a month after the mid-July peak/off-peak split. Continued use after adjustment constitutes acceptance. If your cost model has a DeepSeek number in it, that number is stale.
Each link below shares sources, entities, or timing with this story.
DeepSeek-V4-Flash-0731 landed July 31 under MIT with a DSpark speculative-decoding module attached. Terminal Bench 2.1: 82.7. Toolathlon-Verified: 70.3. DSBench-FullStack: 68.7. DeepSWE: 54.4. NL2Repo: 54.2. The model card claims it beats DeepSeek-V4-Pro (Preview) "despite its...
GPT-5.6 Luna went to $0.20 input / $1.20 output per million tokens on July 30. That's an 80% cut. Terra dropped 20%. Luna's input now undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's $1 input. Simon Willison covered the announcement and...
Accio open-sourced it August 2, now 1,042 stars: 107 tasks across 14 containerized mock services replicating Alibaba, Shopify, and FreightOS (GitHub). Task mix is 53 CLI, 28 browser, 16 file ops, 10 API/MCP, split 65 text-only / 20 browser-text / 22 vision-requiring. Twelve mo...
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
Hugging Face published its Summer 2026 State of Open Models report on August 14, and one statistic in it went almost entirely unremarked in the coverage. By July 2026, agents rather than humans became the Hub's primary users. Claude Code alone accounted for 44.4% of all agent...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.