Fetching from the wire…
Infra2026-06-25 · source-backed
NVIDIA and AWS announced June 23 that NVIDIA's cuVS library now powers GPU-accelerated vector indexing as the default in Amazon OpenSearch Serverless, claiming up to 10x faster index builds at roughly a quarter the cost versus CPU-only, making billion-scale vector DBs buildable in under an hour (NVIDIA Blog). The same announcement debuts EC2 G7 instances on RTX PRO 4500 Blackwell. For anyone running RAG at scale, index build time has been a real operational tax. This attacks it directly, and "default" means you get it without re-architecting. The retrieval layer keeps getting cheaper, which keeps eroding the case for boutique vector-DB vendors.
Each link below shares sources, entities, or timing with this story.
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
NVIDIA's Blackwell successor is in production ahead of schedule. The NVL72 rack (72 GPUs) delivers 3.6 exaFLOPS for inference, with 288GB HBM4 per GPU. NVIDIA claims 10x lower cost-per-token versus Blackwell. The Rubin CPX variant — purpose-built for million-token inference —...
Nvidia set new MLPerf Inference v6.0 records on April 2 using four GB300 NVL72 systems (288 Blackwell Ultra GPUs) interconnected via Quantum-X800 InfiniBand. The headline number: 2.49 million tokens per second on DeepSeek-R1 in offline mode. That's the largest GPU configuratio...
Six co-designed chips, supply chain twice the size of Grace Blackwell, with AWS, Google Cloud, Microsoft, and OCI deploying instances in H2 2026 (NVIDIA). If inference really drops 10x, the economics of always-on agents change at the root. The cost crisis in story one is partl...
The June 30 update adds open-weight models for speech and multimodal retrieval, open datasets, and libraries aimed at agentic AI, while Palantir launched an engine to run Nemotron in air-gapped US government environments, and Nemotron 3 Ultra has day-0 vLLM support. This is th...
$12,930,300,000. Not "about $13 billion." That exact figure is in Jensen Huang's September 3 post, and the precision is the first sign this is a signed agreement rather than the leak-stage reporting we got in late August. NVIDIA's blog frames the platform in the numbers that m...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.