Fetching from the wire…
Public story · 2026-08-05 · high
The proof needs no unproven assumptions, unlike most complexity separations that lean on things like P not equaling NP.
Why now: The paper posted to arXiv in August 2026.
Constant-depth quantum circuits can do two things no constant-depth transformer or diffusion language model can match, an arXiv paper proves unconditionally.
That caps a scaling story a lot of people treat as infinite. One proof shows a function built from an O(log log n)-depth quantum circuit plus one AND gate. Any constant-depth transformer computing it needs width that grows as n^Ω(1). More width, not more layers, is the only way a fixed-depth model keeps up.
The other proof targets diffusion language models directly. It shows a distribution that constant-depth quantum circuits can sample. No constant-round diffusion model can match it within constant distance, not even with shallow denoising schedules, sublinear chain-of-thought, or token revision and remasking.
Both results are unconditional. Rare, for this field. Most separations in complexity theory lean on unproven assumptions, like P not equaling NP. These don't, which is what makes the paper stand out even before you get to what it's separating.
The paper doesn't say what happens once transformer or diffusion depth is allowed to grow with input size instead of staying fixed. That's a different regime with different rules, and it's the regime real production models arguably live in.
My take: depth, not scale, is the wall these proofs describe. Feeding a fixed-depth model more chain-of-thought tokens or more remasking passes doesn't get around either result, both hold with those add-ons already in play. Worth watching whether follow-up work pushes these separations into the depth-grows-with-input regime, where most deployed models actually sit. It posted to arXiv in August 2026.
Each link below shares sources, entities, or timing with this story.
Two papers land the same week: DreamReasoner-8B uses block-size curriculum learning to scale parallel block-wise denoising for long chain-of-thought (arXiv 2606.19257), and Diffusion-Proof applies diffusion-style generation to formal theorem proving (arXiv 2606.19315). The cas...
Diffusion LMs decode many tokens per step but pay to interact with all suffix tokens every step, and existing fixes just keep a local window while re-initializing suffix tokens identically each timestep (arXiv 2608.23167). This method splits the suffix into local, middle and t...
Orr Paradise, Oliver Richardson, Yoshua Bengio and Shafi Goldwasser construct an interactive PCP protocol where a polynomial-time verifier certifies approximate consistency of predictions specified by circuits, even though those circuits implicitly define exponentially many cl...
arXiv 2608.05424 shows ImageNet- and LAION-scale pretrained encoders pick up metadata traces tied to camera and image-processing properties. Deliberately injecting metadata-semantics correlations during pretraining produces systematically higher metadata sensitivity and larger...
Four preregistered studies (1,584 multi-agent simulations, 16 languages, 3 model families) prove that alignment interventions reducing harmful outputs in English actively amplify them in Japanese and 14 other languages. Alignment-induced dissociation correlates with Power Dist...
Ant Group's inclusionAI released a 100B non-embedding-parameter MoE diffusion model adding Levenshtein Editing, explicit DELETE and INSERT control tokens, so diffusion decoding can restructure a sequence and open insertion slots during parallel generation instead of only unmas...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.