Fetching from the wire…
Top 5 · 2026-04-03 · source-backed
Nvidia set new MLPerf Inference v6.0 records on April 2 using four GB300 NVL72 systems (288 Blackwell Ultra GPUs) interconnected via Quantum-X800 InfiniBand. The headline number: 2.49 million tokens per second on DeepSeek-R1 in offline mode. That's the largest GPU configuration ever submitted to any MLPerf benchmark.
The number that matters more for builders: 250K tokens/sec on the interactive benchmark at 30 cents per million tokens generated. That's a 2.77x speedup over the prior-generation GB200 NVL72.
Nvidia was the sole platform to submit across all new tests, including Qwen3-VL-235B and text-to-video generation. Nobody else could even run the full suite.
I keep coming back to the 30 cents number. Right now, if you're calling Claude or GPT APIs at scale, you're paying somewhere between $3 and $75 per million output tokens depending on the model. Self-hosted inference on Blackwell Ultra at $0.30/M is an order of magnitude cheaper than most API pricing. Yes, the upfront hardware cost is enormous. Yes, you need the expertise to run it. But for companies processing millions of requests daily, the build-vs-buy math just shifted hard.
This also matters for the open model ecosystem. vLLM just crossed 75K stars with expanded Blackwell support. NVIDIA is optimizing Gemma 4 for deployment across RTX to DGX Spark to Jetson. The inference stack is maturing fast enough that "run your own models" is becoming a real option for mid-size companies, not just hyperscalers.
If you're planning inference infrastructure for the next 12 months, these benchmarks are your baseline. The 30-cent floor reshapes every cost model I've seen.
Each link below shares sources, entities, or timing with this story.
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
Google released Gemma 4 on April 2 with four model variants: E2B, E4B, 26B MoE, and 31B Dense. The license change is the first thing worth noting. Every previous Gemma had restrictions that made lawyers nervous. Gemma 4 is Apache 2.0. Full stop. Use it in any product, any way...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
The comparison is against GB300 NVL72, with 35x lower cost per million tokens, measured on the SemiAnalysis AgentX benchmark using real recorded agentic coding sessions with context growth, tool calls and sub-agent spawning preserved (NVIDIA). DeepSeek V4 Pro and Qwen3.5 were...
NVIDIA and AWS announced June 23 that NVIDIA's cuVS library now powers GPU-accelerated vector indexing as the default in Amazon OpenSearch Serverless, claiming up to 10x faster index builds at roughly a quarter the cost versus CPU-only, making billion-scale vector DBs buildabl...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.