Fetching from the wire…
Public story · 2026-08-25 · high
On a fixed $100 budget, GLM-5.3 finished 17 DeepSWE tasks to Fable 5's three, and two more benchmarks show cheap models winning on cost per completed task.
Why now: AWS put the GPT-5.6 family into Kiro on August 24, adding a fourth cost-per-task comparison alongside the DeepSWE and Ed-o-meter numbers.
GLM-5.3 finished 17 DeepSWE tasks on a fixed $100 budget; Fable 5 finished three, per Together AI's test, reported by Latent Space AINews. Their first-try accuracy was close, so the five-to-one gap came entirely from how many attempts $100 buys. For anyone paying frontier prices on agent loops that retry after failure, that's the number that should set the budget.
The same report puts GPT-5.6 Sol Max at 72.7% on DeepSWE v1.1 for $6.47 per task, against Fable 5 Max's 69.7% for $21.63. Three points of accuracy for 3.3 times the price.
Ed Yau's Ed-o-meter runs seventeen models through the same 28 tasks with identical prompts and deterministic grading. It puts GLM-5.3 at a 100% pass rate, a 9.3 rubric score, and $0.28 per task. GPT-5.5 answers faster, 13.2 seconds to first token against 16.3, but costs about five times as much. Yau's caveats matter here. One trial per task, wide statistical intervals, and rubric scores Fable 5 assigned itself against saved answer text.
Kiro adds a fourth data point. AWS put the GPT-5.6 family into Kiro on August 24 with a three-tier credit multiplier, Sol at 2.4x, Terra at 1.2x, Luna at 0.6x. Terra costs about 82% less per successful Terminal-Bench 2.1 task while scoring 74.6 on the Coding Agent Index, against Opus 4.8's 72.5, according to the Kiro blog.
If your agent loop retries after a failure, and most do, the number that decides your bill is completed tasks per dollar, not first-attempt accuracy. That only holds if whatever checks task completion is honest about which attempt worked. Feed retries to a weak verifier and a cheap model just produces plausible garbage faster.
Each link below shares sources, entities, or timing with this story.
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill. SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-...
Everyone kept score wrong. When OpenAI shipped GPT-5.6 (the Sol flagship plus Terra and Luna) to GA on July 9, then xAI put out Grok 4.5, Meta dropped Muse Spark 1.1, and Cognition shipped SWE-1.7, the reflex was to ask who won the benchmark. Wrong question. On the Artificial...
You can't sign up for the best coding model OpenAI has ever built. You have to be approved. By the federal government. One customer at a time. OpenAI previewed GPT-5.6 'Sol' on June 26, and the capability story is real: it's a three-model family (Sol the flagship at $5/$30 per...
This is the other half of the Fable 5 story, so read them together. While the best coding model in the world is uncallable, an open-weight one quietly posted frontier-adjacent numbers. Per Tom's Hardware, independent benchmarks for the MIT-licensed GLM-5.2 (744B params, 40B ac...
OpenAI launched the GPT-5.6 family on July 14: Sol (flagship), Terra (cost-optimized), and Luna (fast tier), live across ChatGPT, Codex, and the API the same day after a US-government-requested delay for security review. The numbers are loud. Sol scored 53.6 on Agents' Last Ex...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.