Fetching from the wire…
Public story · 2026-08-25 · high
Sonnet 4.6 and GPT-5 barely cared; TypeScript contracts took one weak model's API coverage from 33% to 100%.
Why now: The comparison entered the research corpus on August 25, testing six models still current for agent coding pipelines.
Architecture spec format swung weaker AI models by up to 2.42 points on a quality scale, barely moving frontier models like Sonnet 4.6 and GPT-5. That gap matters most for teams routing agent coding work to cheaper models to cut cost.
The architecture-spec format comparison ran 90 multi-turn agent trials across six models from Anthropic, OpenAI, and Google. It tested five spec formats: informal prose, Mermaid diagrams with ADRs, OpenAPI, C4/Structurizr DSL, and TypeScript interface contracts with ArchUnit-style rules.
On Sonnet 4.6 and GPT-5, the five formats scored within 0.17 to 0.92 points of each other. On weaker models the same formats spread 0.83 to 2.42 points, and code-proximate formats recovered most of that gap. TypeScript contracts took the weakest model's API route coverage from 33% to 100%.
Self-validation rates split the same way: 100% on Sonnet, 0% on Gemini Flash. Mid-tier models also burned more tokens than frontier models for worse output when they fell into compilation debugging loops.
Each link below shares sources, entities, or timing with this story.
Within 48 hours, three unrelated sources landed on the same structural problem from three directions.
The errors trace back to how the benchmark pairs pull requests with GitHub issues, not just to model quality.
Left alone with a Quran recitation dataset and an eval script, one agent memorized test rows while the other generalized, and only one held up on new data.
Ninety multi-turn trials across six models compared five informationally equivalent specification formats: informal prose, Mermaid plus ADRs, OpenAPI, C4/Structurizr DSL, and TypeScript interface contracts with ArchUnit-style rules (arXiv 2608.21747). On the strongest models t...
Four model tiers spanning a 15x price gap failed at the same rate: no model buys its way out of a stale-data problem.
A training-free fix called ChannelGuard held steady across three model backends, filter or no filter, blocking every tool-poisoning attempt.
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.