Fetching from the wire…
Top 5 · 2026-08-02 · source-backed
The number that reframes everything isn't ten. It's two thousand.
OpenAI published "Ten advances in mathematics and theoretical computer science" on August 1, claiming an internal version of Astra produced new results on ten problems that had seen no progress on the main result for at least a decade. High-dimensional sphere packing. Connes's rigidity conjecture. Ehrhart's volume conjecture. Multicolor Ramsey numbers. Quantum parallel repetition. Arithmetic circuit complexity. The closest vector problem. Non-sofic groups. Binary and spherical codes. Extremal number conjectures.
They shipped receipts. The openai/ten-proofs repo is Apache-2.0, Lean 4.32.0, sitting around 275 stars, with machine-checkable certificates alongside an LLM-written PDF reconstructing the derivations from unpublished reasoning traces. Machine-checkable is the part that separates this from every previous "AI solved math" claim. You don't have to trust the narrative. You can run the kernel.
Then Noam Brown posted the cost: all ten proofs, combined, ran under $2,000 at GPT-5.6 Sol API prices. He also killed the obvious follow-up question directly. "Sadly no Millennium Prize problems (yet)." They tried other major problems and failed. They didn't spend heavily per problem, and test-time compute could be pushed much further.
That last clause is the actual story. If frontier mathematical discovery costs $200 a problem and nobody has pushed the compute knob hard, then this isn't a capability that arrives with the next model. It's a dial that's already installed and turned down. That's a completely different planning assumption than "wait for GPT-6."
The mathematicians are not celebrating. Terence Tao described GPT-5.6 Pro solving problems he'd personally spent time on as "very strange and not particularly pleasant." Timothy Gowers, a Fields medalist, warned about the "possible destruction of mathematical culture" as the literature expands while human understanding thins. Queen Mary's Abhishek Saha landed somewhere more useful: frontier models are "at least as good as a solid and indefatigable PhD student," and he now plays "conductor, rather than doubling up as the whole orchestra." Kirwin Hampshire called the reassuring framing "a well-muffled scream," and his July essay resurfaced on r/singularity today at 356 upvotes and 513 comments, a 1.44 comment-to-score ratio that means people are arguing, not nodding.
Willison flagged the calibration problem nobody else did: OpenAI discloses cost per success but not the denominator. How many problems did Astra attempt? How many did it fail? Ten wins out of ten attempts and ten wins out of four hundred are the same press release and completely different technologies.
I'd add a second gap. Thomas Bloom at Manchester called it "big news" while noting nobody has had time to referee these at the depth these conjectures normally get. A Lean certificate proves the theorem follows from the axioms. It doesn't tell you how much human problem-shaping happened before the run, which is exactly what Tao's "tireless literature-scanning assistant" framing is pointing at.
What to do with this: stop treating frontier reasoning as a fixed capability tier you shop for. If the cost curve on hard-problem solving is this low and this unexplored, the question for your own work isn't "which model" but "how much compute am I willing to spend on one problem." Most of us have never asked. And when a rebuttal PDF claiming the Connes disproof was invalid got shredded on HN within hours, the same author has a claimed Riemann Hypothesis proof, and commenters flagged the rebuttal itself as likely AI-generated, the real open question surfaced: when validation costs this much effort, how does anything get adjudicated at all?
Each link below shares sources, entities, or timing with this story.
Quanta's August 3 piece tallies the assault: OpenAI found a counterexample to Erdős's 1946 unit distance conjecture on May 20, then Astra produced 10 further advances. Google DeepMind evaluated 700 open conjectures in January, solving four and recovering nine forgotten solutio...
GPT-5.6 Luna went to $0.20 input / $1.20 output per million tokens on July 30. That's an 80% cut. Terra dropped 20%. Luna's input now undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's $1 input. Simon Willison covered the announcement and...
The company published "Pacing model development in an era of cyber-critical capabilities" on August 19, disclosing the pause on its latest deployment-bound models while it hardened and red-teamed research environments. The trigger was an unreleased model, Astra, plus a July in...
OpenAI disclosed August 7 that internal evaluations of the upcoming Astra model show agentic coding and cybersecurity performance strong enough that it can no longer rule out the Critical cybersecurity level in its Preparedness Framework, a first. Every prior frontier model in...
Promptwatch's tracking shows the share of ChatGPT search queries using site: sat at 0.3-0.5% for weeks, dipped to 0.15% on August 3-5, then jumped to 16-17% on August 8, two days after OpenAI said it was making GPT-5.6 Sol "more reliable with facts." Simon Willison Willison co...
His August 2 post concedes the result is real and attacks the inference as a fallacy of composition: success on one form of fancy cognition doesn't mean success on all forms is imminent. His sharpest technical objection is that math is uniquely favorable because it "allows for...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.