Fetching from the wire…
Public story · 2026-03-14 · source-backed
Jensen Huang delivers his GTC 2026 keynote March 16 at SAP Center (11am PT), promising "a few new chips the world has never seen before" with focus on agentic-optimized CPUs and a CPU-only rack — representing a strategic shift from GPU-centric toward heterogeneous inference hardware. Expected: Rubin GPUs with 288GB HBM4, 5x Blackwell throughput, and OpenClaw platform details. 30,000+ attendees.
Each link below shares sources, entities, or timing with this story.
GTC 2026 runs March 16-19 in San Jose. Expected: Vera Rubin architecture deep-dive (VR200 NVL72 delivering 3.3x inference performance vs Blackwell Ultra), possible Feynman architecture early samples (TSMC A16 1.6nm with silicon photonics — optical signals replacing electrical...
NVIDIA and AWS announced June 23 that NVIDIA's cuVS library now powers GPU-accelerated vector indexing as the default in Amazon OpenSearch Serverless, claiming up to 10x faster index builds at roughly a quarter the cost versus CPU-only, making billion-scale vector DBs buildabl...
NVIDIA's Blackwell successor is in production ahead of schedule. The NVL72 rack (72 GPUs) delivers 3.6 exaFLOPS for inference, with 288GB HBM4 per GPU. NVIDIA claims 10x lower cost-per-token versus Blackwell. The Rubin CPX variant — purpose-built for million-token inference —...
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
Blaizzy/nativ (1,163 stars, Swift, MIT, macOS 26+) comes from the mlx-vlm author and bundles that server into a SwiftUI app that discovers MLX models already in your HF cache. It exposes OpenAI-compatible chat, Responses, image, audio and model endpoints plus Anthropic Message...
Microsoft open-sourced Foundry-Local, and I think it's the most underrated release of the week. One SDK for chat and audio with automatic hardware acceleration (NPU > GPU > CPU), self-contained with no external dependencies, and the API surface is identical to Azure AI Foundry...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.