Fetching from the wire…
Infra2026-06-08 · source-backed
At his June 1 GTC Taipei keynote, Huang reframed computing as stating intent to an AI that writes code or calls tools, and revealed Vera Rubin, the next platform after Blackwell, purpose-built for agentic training and inference. When the hardware roadmap is explicitly indexed to agent workloads, that's where compute demand is heading. The whole stack from silicon up is reorganizing around agents, not chat.
Each link below shares sources, entities, or timing with this story.
On the August 26 earnings call Huang said "AI has reached its inflection point. It's doing useful work. Its tokens are productive and profitable. Now, compute is revenue," and dismissed the AGI debate as "kind of senseless" in favor of whether AI does profitable work (CNBC). N...
GTC 2026 runs March 16-19 in San Jose. Expected: Vera Rubin architecture deep-dive (VR200 NVL72 delivering 3.3x inference performance vs Blackwell Ultra), possible Feynman architecture early samples (TSMC A16 1.6nm with silicon photonics — optical signals replacing electrical...
Six co-designed chips, supply chain twice the size of Grace Blackwell, with AWS, Google Cloud, Microsoft, and OCI deploying instances in H2 2026 (NVIDIA). If inference really drops 10x, the economics of always-on agents change at the root. The cost crisis in story one is partl...
TechCrunch's August 29 piece frames Nvidia's durable advantage as system-level, built around Vera Rubin pairing the Rubin GPU with the Vera CPU, a Groq 3 LPX inference accelerator, and storage and networking racks. VP of storage technology Jason Hardy is quoted claiming "upwar...
In a July 17 post, NVIDIA pushes "intelligence per dollar" as the agentic-era metric, arguing post-training rather than pretraining is now the central workload because agents need continuous improvement cycles. The claim: a 10-trillion-parameter MoE on 100 trillion tokens in o...
NVIDIA's Blackwell successor is in production ahead of schedule. The NVL72 rack (72 GPUs) delivers 3.6 exaFLOPS for inference, with 288GB HBM4 per GPU. NVIDIA claims 10x lower cost-per-token versus Blackwell. The Rubin CPX variant — purpose-built for million-token inference —...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.