Fetching from the wire…
Public story · 2026-08-25 · high
The Vera Rubin NVL72 numbers come from recorded coding sessions with tool calls and sub-agent spawning intact, not single-turn benchmarks.
Why now: The comparison surfaced in coverage dated August 25, with the numbers still unverified by outside review.
Vera Rubin NVL72 reaches 30x the throughput per megawatt of the current GB300 NVL72, per Nvidia's Vera Rubin efficiency post. The company also claims a 35x drop in cost per million tokens. That number matters most for teams sizing agent deployments, where a session's context grows across tool calls and sub-agent spawns before finishing a task.
Nvidia measured both figures on SemiAnalysis's AgentX benchmark, built from recorded agentic coding sessions rather than single-turn prompts. DeepSeek V4 Pro and Qwen3.5 were among the models tested. AgentX scores a full session. It counts context growth, tool calls, and sub-agent spawning, not a one-shot question and answer.
The results are vendor-reported and still waiting on SemiAnalysis's own review. There's no GA date for Vera Rubin NVL72, and Nvidia's post doesn't say when that review will publish.
The benchmark choice signals a shift in how chip vendors market inference economics, from single-turn token counts to full agent sessions. Watch for SemiAnalysis to publish its own AgentX numbers on the same hardware, independent of Nvidia's release, before treating 30x as settled.
Each link below shares sources, entities, or timing with this story.
Nvidia set new MLPerf Inference v6.0 records on April 2 using four GB300 NVL72 systems (288 Blackwell Ultra GPUs) interconnected via Quantum-X800 InfiniBand. The headline number: 2.49 million tokens per second on DeepSeek-R1 in offline mode. That's the largest GPU configuratio...
DeepClaude hit 470 points on Hacker News. It swaps Claude Code's API backend to DeepSeek V4 Pro while preserving the full agent loop: file editing, bash execution, git tooling, the whole workflow. DeepSeek V4 Pro scores 96.4% on LiveCodeBench at a fraction of Anthropic's prici...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
NVIDIA claims up to 4x faster output and 30% faster agentic task completion versus comparable models, aimed explicitly at high-volume narrow work: code review, tool use, security monitoring, billing triage (NVIDIA). They also published Nemotron-RL-Agentic-Terminal-Pivot, an ag...
NVIDIA's GEAR Lab, with CMU and UC Berkeley, released a closed-loop system where coding agents reset physical scenes, run hardware trials, verify outcomes, and rewrite code until a policy works. Jim Fan calls it "AutoResearch in the physical world." Agent teams hit 99% pass@8...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.