Fetching from the wire…
Public story · 2026-08-25 · high
The interactive inference accelerator claims 3,400 output tokens a second, and Nebius is the first cloud to run it.
Why now: Nvidia disclosed the production milestone on August 24 at Hot Chips, its venue for staking inference-speed claims.
Nvidia put Groq 3 LPX into full production on August 24, claiming 3,400 output tokens per second, according to Nvidia's announcement at Hot Chips.
LPX targets the token-generation phase, the step that decides how responsive an agent feels. That phase is what stalls while a loop waits on a multi-step tool call. Nvidia calls it an interactive inference accelerator and says it runs 4x faster on agent responsiveness than the nearest alternative.
The 3,400-tokens-per-second figure comes from running Gemma 4 31B at a 100,000-token context. LPX extends the Vera Rubin NVL72 platform.
Nebius is the first cloud provider to put it into production. Nvidia hasn't disclosed pricing, so what that speed costs at scale is still unknown.
That figure is Nvidia's own benchmark, run on Nvidia's own comparison set. It stays a marketing number until an independent test reproduces the 4x agent-responsiveness claim against a workload nobody at Nvidia picked. Watch for Nebius or a third party to publish that number once LPX is live outside Nvidia's demos.
Each link below shares sources, entities, or timing with this story.
NVIDIA claims 1.8x faster task completion and twice the efficiency against traditional x86, with Vera Rubin NVL72 racking 72 Rubin GPUs and 36 Vera CPUs alongside ConnectX-9 SuperNICs and BlueField-4 DPUs. (NVIDIA) Architecture detail months before shipping, two days ahead of...
440 stars, created August 11, covering 12 sections from model APIs through structured output, RAG, evals, agent loops, LoRA versus fine-tuning, security, LLMOps and serving, plus three case studies and a capstone (GitHub). The stance is that you write the agent loop, RAG and e...
Six co-designed chips, supply chain twice the size of Grace Blackwell, with AWS, Google Cloud, Microsoft, and OCI deploying instances in H2 2026 (NVIDIA). If inference really drops 10x, the economics of always-on agents change at the root. The cost crisis in story one is partl...
Three moves, two days, no coordination between them. August 10–11: GitHub shipped Ollama as a BYOK provider inside Copilot for JetBrains (GitHub Changelog). Unsloth released Unsloth Desktop with a command literally named unsloth start claude, which points Claude Code and Codex...
Portable Computer launched August 26, running the orchestrator LLM, subagent LLM, planner, tool router, scheduler and local search index locally, with local work consuming no billing credits and each cloud escalation requiring separate approval (VentureBeat). Launch platform i...
NVHBM, announced August 26, relocates NVIDIA's custom memory controller from the compute chip into the HBM base die, claiming up to 30% more bandwidth than standard HBM4E, 15% lower HBM power, and up to 25% more freed area on the XPU compute die (NVIDIA). Next-generation Train...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.