Fetching from the wire…
Research2026-06-18 · source-backed
Two papers land the same week: DreamReasoner-8B uses block-size curriculum learning to scale parallel block-wise denoising for long chain-of-thought (arXiv 2606.19257), and Diffusion-Proof applies diffusion-style generation to formal theorem proving (arXiv 2606.19315). The case that diffusion LMs can compete with autoregressive models on structured reasoning, while decoding in parallel, is getting harder to wave off as a curiosity.
Each link below shares sources, entities, or timing with this story.
arXiv 2608.03962 proves two unconditional separations. First, a distribution sampleable by constant-depth QNC^0 circuits that no constant-round diffusion language model with shallow scheduling and denoising can sample within constant distance, even given sublinear chain-of-tho...
Per Hugging Face Daily Papers, PerceptionDLM processes multiple visual regions in parallel using diffusion-based language models instead of autoregressive decoding. It's another data point that diffusion-LM architectures are getting real traction in vision-language, where para...
Co-led by Lightspeed and Diffusion with Sapphire and Whale Rock joining, up from $11B in March, on ARR past $350-400M with 80% of Am Law 100 firms as customers. Capital is earmarked for proprietary models rather than wrapping frontier ones. Bloomberg A vertical app company buy...
Amid a week of pricing and commerce stories, here's hard tech you can actually download. Google released DiffusionGemma on June 10, a 26B-parameter Mixture-of-Experts model (3.8B active) that generates text by diffusion instead of left-to-right decoding. The architecture is th...
Think in Diffusion, Talk in Autoregression: drafts tokens in parallelized diffusion then outputs autoregressively in a single forward pass. The draft model is the base model itself. First architecture to close the quality gap with pure AR at nearly 6x throughput. Source ---
Diffusion LMs decode many tokens per step but pay to interact with all suffix tokens every step, and existing fixes just keep a local window while re-initializing suffix tokens identically each timestep (arXiv 2608.23167). This method splits the suffix into local, middle and t...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.