Fetching from the wire…
Infra2026-08-06 · source-backed
Announced at FMS, the cuFile APIs let GPUs read and write storage directly in microseconds rather than routing through CPUs, and SCADA (scaled, accelerated data access) lets massively parallel GPUs pull only application-necessary data into high-bandwidth memory. NVIDIA also showed the Vera CPU in the Vera BlueField-4 STX storage processor delivering up to 3.21x higher throughput than an x86 CPU on a two-stage compression-and-encryption pipeline. 40+ storage and flash vendors are in the Storage-Next initiative. The driver: KV-cache and long-context working sets outgrowing system memory.
Each link below shares sources, entities, or timing with this story.
NVHBM, announced August 26, relocates NVIDIA's custom memory controller from the compute chip into the HBM base die, claiming up to 30% more bandwidth than standard HBM4E, 15% lower HBM power, and up to 25% more freed area on the XPU compute die (NVIDIA). Next-generation Train...
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
TechCrunch's August 29 piece frames Nvidia's durable advantage as system-level, built around Vera Rubin pairing the Rubin GPU with the Vera CPU, a Groq 3 LPX inference accelerator, and storage and networking racks. VP of storage technology Jason Hardy is quoted claiming "upwar...
NVIDIA claims 1.8x faster task completion and twice the efficiency against traditional x86, with Vera Rubin NVL72 racking 72 Rubin GPUs and 36 Vera CPUs alongside ConnectX-9 SuperNICs and BlueField-4 DPUs. (NVIDIA) Architecture detail months before shipping, two days ahead of...
Jensen Huang disclosed 100% of NVIDIA uses Claude Code, calling it "the first agentic model." Vera CPU: 88 custom Olympus cores, 1.2 TB/s LPDDR5X, paired with Rubin GPUs at 1.8 TB/s coherent bandwidth. 22,500+ concurrent CPU environments per rack. Dell, HPE, Lenovo, Alibaba, B...
Released in beta at IFA 2026, PAIR distributes inference requests across whichever PCs on a local network have spare capacity, so agentic workflows run in parallel. Supports GeForce RTX 20 Series and newer, RTX PRO workstation GPUs from Turing on, DGX Spark, and Apple M4 or ne...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.