Fetching from the wire…
Top 5 · 2026-08-14 · source-backed
Netlify published an AXIS-framework evaluation on August 14 that I've been thinking about all day. Same task, 11 models, three runs each, scored on functional correctness rather than aesthetics. The task was deliberately boring: a static one-page coffee shop site with hours, address, menu and a photo. Explicitly no CMS.
Average credit cost, cheapest to most expensive: DeepSeek V4 Flash 0731 at 2.4. Kimi K2.7 Code at 19. GLM 5.2 at 27. GPT 5.6 Terra at 39. Gemini 3.1 Pro at 53. Kimi K3 at 102. GPT 5.6 Sol at 141. Claude Sonnet 5 at 143. Claude Opus 5 dead last at 519.
That's a 216x spread on an identical task where all of them produced a working page. Opus 5 spent 216 times what DeepSeek V4 Flash spent to build the same coffee shop website.
I've done this to myself. I have a default in my head that says "use the best model, it's worth it," and for hard architectural work it usually is. For scaffolding a static page it's just setting money on fire. Frontier reasoning models overspend on simple work because they're built to explore, and simple work has nothing to explore.
The fix arrived the same week. LLMRouter from Tao Feng, Jiaxuan You and colleagues at UIUC (arXiv 2608.06867) hit 2,340 HuggingFace upvotes: learned routers outperform the strongest fixed-model baseline by 14.6% relative, and lightweight routers get more competitive as cost constraints tighten. They open-sourced 16+ representative routers and an xRouteBench evaluation platform covering single-turn, multi-turn and personalized routing. The problem and its answer landed within a week of each other.
Urgency on this just went up. DeepSeek raised API prices 50% to more than 1,100% depending on model, token type and time of day, effective 16:00 UTC on August 16. V4-Pro cache-miss input goes from $0.435 to $1.32 per million at peak, with peak windows at 01:00–04:00 and 06:00–10:00 UTC and off-peak at half. If you have V4 in a production loop, you have two days to re-price.
Do this today: pick the three dumbest, highest-volume tasks in your pipeline. Boilerplate generation, file scaffolding, commit message writing. Route them to a cheap model and diff the output against what your expensive model produces. If you can't tell the difference, you just found your margin.
Each link below shares sources, entities, or timing with this story.
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
The leaderboard says first place. The methodology says you should check your own bill. Qwen3.8 Max now ranks first on Artificial Analysis' agentic index, scoring 86.1 on OSWorld-Verified ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0, priced at $2.00/M input and $6.00/M...
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
Writer launched Palmyra X6 on August 13 with a number that should reset how you think about agent COGS: 52% lower average cost, 48% better speed, 10% better quality. The model is a post-training variation of Z.ai's open-source GLM-5.2. A US enterprise SaaS vendor built its fla...
GPT-5.6 Luna went to $0.20 input / $1.20 output per million tokens on July 30. That's an 80% cut. Terra dropped 20%. Luna's input now undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's $1 input. Simon Willison covered the announcement and...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.