Fetching from the wire…
Public story · 2026-07-13 · high
The framework arrives with two Gemini Deep Think papers and no consensus on how to measure an AI's share of a proof.
Why now: The proposal appeared in coverage dated July 13, alongside DeepMind's two new Deep Think papers pushing models from benchmark scores toward open problems.
DeepMind proposed a taxonomy for classifying AI-assisted mathematics by significance and degree of AI contribution, following consultation with the mathematics community, per Google DeepMind. The credit question matters: if a model does 80% of a proof, who gets the paper, and how do you measure that 80%? Nobody has a clean answer yet.
The taxonomy arrived alongside two new papers built on Gemini Deep Think. That's DeepMind's effort to push its models past benchmark scores and into open problems that don't come with a known answer.
Benchmarks just check whether a model lands on the right number. Open problems need someone, human or machine, to convince other mathematicians the proof actually holds. That's a different kind of credit than acing a test.
Software rarely fights over who wrote a function once a model drafts it and a person edits it. Math can't shrug that off the same way, because a proof's value rides on who's vouching for it. I'd watch whether the next Deep Think paper forces the question: a model-authored proof nobody had to touch, credited how?
Each link below shares sources, entities, or timing with this story.
The pilot puts evaluation inside Confidential Space on Google Cloud's Confidential Computing so Gemini Flash Lite's weights stay private from evaluators while evaluator prompts stay private from Google (DeepMind). Partners are the Singapore AI Safety Institute, OpenMined, AVER...
DeepMind's commitment is denominated in tokens and cloud credits rather than cash, an in-kind structure that keeps the spend on Google infrastructure (DeepMind). OpenAI published its own commitment with the Department of Energy and the national labs the same week (OpenAI). Fed...
On June 22, DeepMind took its first-ever equity position in a film studio, $75M in A24, in a multiyear partnership to co-develop AI filmmaking tools on Veo, with A24 directors testing inside live productions. Demis Hassabis framed it as building "directly with" artists. The de...
Google DeepMind published a one-year impact report for AlphaEvolve showing the Gemini-powered coding agent is now production infrastructure, not research. A Borg scheduling heuristic it generated recovers 0.7% of Google's worldwide compute (in production over a year). A circui...
Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber, the last purpose-built to find and patch vulnerabilities and pitched as a cheap alternative to large security-specialized models like Mythos (DeepMind). Splitting a cheap tier into a security SKU is new packa...
A major regulated bank releasing offensive-style security tooling into the open is genuinely unusual, and it gives builders a production-hardened reference rather than a research prototype (Finextra). It landed the same week Google shipped Gemini 3.5 Flash Cyber, a purpose-bui...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.