Fetching from the wire…
Research2026-05-10 · source-backed
NVIDIA released Star Elastic, a post-training method that nests three submodels (30B, 23B, 12B) inside a single Nemotron Nano v3 checkpoint. The technique uses only 160B tokens (360x reduction vs pretraining) and cuts memory for deploying all three from 126.1GB to 58.9GB in BF16. The clever bit: elastic budget control routes "thinking" through the smaller submodel and only uses the full 30B for the final answer. 16% accuracy gain at 1.9x lower latency. Accepted at ICML 2026.
Each link below shares sources, entities, or timing with this story.
Alibaba published a fine-grained MoE with 2.4T total / 95B active, 512 experts, and a 92-layer hybrid full/linear attention backbone. vLLM shipped day-0 support verified on NVIDIA and AMD with ready 4-bit checkpoints (NVFP4 at 1.32 TiB for an 8xB300 node, MXFP4 at 1.45 TiB for...
Ollama cut v0.34.0-rc1 on September 5 at 23:49 UTC, and the headline item changes the shape of the local-versus-hosted decision rather than the performance of either side: Ollama-hosted open models can be selected directly inside ChatGPT Desktop, with setup driven from the Oll...
paddo.dev makes the most contrarian read: the letter's substance isn't openness but paragraph nine, defending distillation as "a widely used technique for model improvement" and urging policymakers against "conflating legitimate model development techniques with misappropriati...
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
3B active parameters, beats Qwen3.5-35B-A3B on AIME 2025 (92.4 vs 91.9), LiveCodeBench v6 (87.2 vs 74.6), and surpasses the larger Nemotron-3-Super-120B. Available on Ollama and HuggingFace under open license. Source
The XpertStation WS300 pairs NVIDIA's GB300 Grace Blackwell Ultra desktop superchip with 252GB HBM3e at 7.1 TB/s plus 496GB LPDDR5X at 396 GB/s, and dual 400GbE. The r/LocalLLaMA thread is mostly people noting the Newegg listing shows out of stock while MSI's own page has a Ge...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.