Fetching from the wire…
Research2026-08-15 · source-backed
RMM is training-free and input-adaptive, selecting informative slices along contraction dimensions under a single retention-ratio knob. Tested from 1B to 70B across discriminative, autoregressive and long-context settings, reduction tolerance often improved with scale. Custom A100 kernels turned theoretical savings into wall-clock gains, especially at long sequences. arXiv 2608.13426
Each link below shares sources, entities, or timing with this story.
arXiv 2607.14530 gets Hyper-Connections past the N=4 wall by sparsely updating only k=4 streams plus temporal feature augmentation, scoring 4.0 points higher on average downstream than prior mHC on an 18B MoE. Vanilla and mHC need 1.50x and 1.19x xHC's compute to hit the same...
MoT asks whether pretraining can decompose into small independently schedulable jobs. It partitions a Transformer into contiguous layer blocks, trains each inside a frozen pretrained aligner scaffold, then recomposes with an optional short end-to-end adaptation pass (arXiv). O...
Jack Clark's Import AI #461 covers Sequent, founded on the view that frontier labs' reactive approaches won't keep pace with superintelligence development. Instead of reacting, it plans a portfolio of "differentiated alignment bets." Worth noting for anyone tracking where inde...
Talent moves are usually gossip. This one's a signal. Per CNBC, Noam Shazeer, co-author of the Transformer paper and Gemini co-lead, announced June 18 he's leaving Google for OpenAI. The detail that makes it remarkable: Google paid roughly $2.7B two years ago to bring him back...
NVIDIA launched Nemotron 3 Super — a 120B total / 12B active parameter hybrid Mamba-Transformer MoE, open, designed specifically for multi-agent workloads, and delivering 5x higher throughput than Nemotron 2 at the same active parameter count (NVIDIA Newsroom). It ships with a...
Together.ai released Mamba-3 (Apache 2.0, ICLR 2026) — an SSM achieving ~4% better language modeling than the Transformer baseline while running up to 7x faster on long sequences. Key innovations: Exponential-Trapezoidal Discretization, Complex-Valued SSMs with the RoPE Trick,...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.