Fetching from the wire…
Public story · 2026-07-30 · high
JuliaHub's July 30 test found a 0.366-point harness gap, more than double the 0.162 gap between the best and worst model.
Why now: JuliaHub's evaluation and OpenAI's harness settings post both landed July 30, alongside a July 29 study tying agent security complaints to configuration, not model weights.
JuliaHub ran four frontier models through five sealed modeling and simulation problems on July 30. The harness gap beat the model gap by more than double: swapping models produced a 0.162-point spread, swapping the harness produced 0.366. For teams building or shipping agents, that means the scaffolding wrapped around a model matters more than which model you pick.
The hardest problem was a full NASA HL-20 flight vehicle with six-degree-of-freedom dynamics. Claude Fable 5 took the top score at 0.889 and swept all twelve trials on the four core problems. GPT-5.6-Sol followed at 0.814, then GPT-5.6-Terra at 0.786, then GPT-5.6-Luna at 0.727.
OpenAI backed this up the same day. Turning on retained reasoning and compaction in the Responses API took GPT-5.6 Sol from 13.3% to 38.3% on the ARC-AGI-3 public set. Output tokens dropped sixfold, same weights.
Under the official evaluation harness, the same model scored as low as 7.8%. Its private chain of thought got discarded after every move, forcing it to rebuild the game state from scratch each turn. Two settings, nearly a 3x swing.
A July 29 arXiv study adds a third data point. Researchers mined 1.1 million Reddit posts across 29 subreddits and isolated 446 threads on security and privacy problems in Cursor, Copilot, and Codex. After reading over 6,000 comments, most reported issues traced to system-level choices, data access scope, unchecked autonomous actions, not the model underneath.
In my own work, every time output quality drops, my first instinct is to blame the model. It's almost never the model. Usually it's a tool grant revoked in a refactor, or a context budget that quietly truncated.
Before you migrate models, check what your setup can actually reach versus what your prompts assume it can reach. Next time you read a leaderboard, ask whose configuration produced the number.
Each link below shares sources, entities, or timing with this story.
GPT-5.6 Luna went to $0.20 input / $1.20 output per million tokens on July 30. That's an 80% cut. Terra dropped 20%. Luna's input now undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's $1 input. Simon Willison covered the announcement and...
Two thirds. Not two thirds of a contrived jailbreak set. Two thirds of realistic malicious issue requests, against the exact three tools most of the people reading this run daily. Ankur Singh, Jinqiu Yang, and Tse-Hsun Chen built IssueTrojanBench across four attack categories...
OpenAI launched the GPT-5.6 family on July 14: Sol (flagship), Terra (cost-optimized), and Luna (fast tier), live across ChatGPT, Codex, and the API the same day after a US-government-requested delay for security review. The numbers are loud. Sol scored 53.6 on Agents' Last Ex...
July 9, across VS Code, Visual Studio, Copilot CLI, the cloud agent, github.com, GitHub Mobile, JetBrains, Xcode, and Eclipse. Sol is the high-reasoning tier at $5/1M in, $30/1M out, gated to Pro+/Max/Business/Enterprise. Terra is the balanced default at $2.50/$15. Luna is fas...
Here's the experiment: a team of cooperating agents rebuilds SQLite in Rust from scratch, using only the 835-page manual. No source code. No test suites. No internet. Then it has to pass a held-out sqllogictest suite. It worked. Cursor published the research (Wilson Lin, July...
Alongside the July 29 launch of ChatGPT for Academic Researchers, free GPT-5.6 Sol Pro for 10,000 researchers this summer scaling to 100,000 through 2027 backed by over $250 million, OpenAI disclosed efficiency work on the harness underlying Codex and ChatGPT Work: 54% fewer o...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.