Fetching from the wire…
Public story · 2026-03-19 · source-backed
Together.ai released Mamba-3 (Apache 2.0, ICLR 2026) — an SSM achieving ~4% better language modeling than the Transformer baseline while running up to 7x faster on long sequences. Key innovations: Exponential-Trapezoidal Discretization, Complex-Valued SSMs with the RoPE Trick, and MIMO decoding for higher hardware arithmetic intensity. The architecture is explicitly inference-first, targeting agentic workloads where inference — not training — is the bottleneck. Together.ai
Each link below shares sources, entities, or timing with this story.
Together AI published Mamba-3 as an ICLR 2026 paper — new SSM with complex-valued state transitions beating Mamba-2, Gated DeltaNet, and even Llama-3.2-1B on prefill+decode latency across all sequence lengths. The MIMO variant adds +1.2 points average downstream accuracy on to...
Google's MoE model jumped from 6.6% to 86.4% on τ2-bench Retail for tool use. Math +330%, coding +175%. Runs at ~150 tok/s on consumer GPUs. Apache 2.0. The efficiency story here is the real news: 3.8B active parameters achieving Arena AI #6.
NVIDIA debuted the Nemotron 3 family for agentic AI. Nano (30B params, 3B active via MoE) delivers 78% HumanEval, native 1M-token context, and 4x higher throughput than Nemotron 2 via hybrid Mamba-Transformer architecture. Critical developer angle: NVIDIA open-sources NeMo Gym...
RMM is training-free and input-adaptive, selecting informative slices along contraction dimensions under a single retention-ratio knob. Tested from 1B to 70B across discriminative, autoregressive and long-context settings, reduction tolerance often improved with scale. Custom...
Agentic Resource Discovery went up August 24 at agenticresourcediscovery.org under Apache 2.0, contributed to but not authored by AWS (AWS ML Blog). It lets agent, tool and skill registries federate across clouds, on-prem and SaaS without bilateral connectors. AWS Agent Regist...
Cohere launched North Mini Code on June 9 under Apache 2.0, its first developer-focused model. The shape is the pitch: 30B parameters, mixture-of-experts, only ~3B active, and it runs on a single H100. It scores 33.4 on the Artificial Analysis Coding Index, competes on SWE-Ben...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.