Fetching from the wire…
Infra2026-08-15 · source-backed
The French startup argues the belief that GPUs suit agentic workflows poorly is a misconception, and says its Kog Inference Engine reaches up to 30x decoding speedups on standard NVIDIA and AMD datacenter GPUs with no new hardware. The approach is hardware-aware optimization of memory bandwidth use, not a new architecture. TechCrunch Treat 30x as a vendor claim on selected workloads until someone benchmarks it independently. The thesis directly contradicts the custom-inference-silicon narrative, which is why it's worth watching either way.
Each link below shares sources, entities, or timing with this story.
AMD unveiled its first rack-scale system to directly contest Nvidia at the rack level, with engineering samples in H2 2026 and mass production targeted Q2 2027. Microsoft joins Meta, OpenAI and Oracle as customers; Meta plans 1 gigawatt of Helios racks by year-end against a lo...
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
A model-routing API covering OpenAI, Anthropic, DeepSeek, Moonshot, Minimax, Nvidia, xAI, and Z.ai, with strategies including preferring flex usage tiers, specifying up to three performance benchmarks for automatic selection, or sending only complex queries to expensive models...
NVIDIA's Blackwell successor is in production ahead of schedule. The NVL72 rack (72 GPUs) delivers 3.6 exaFLOPS for inference, with 288GB HBM4 per GPU. NVIDIA claims 10x lower cost-per-token versus Blackwell. The Rubin CPX variant — purpose-built for million-token inference —...
June local-inference benchmarks across the 128GB class put NVIDIA's DGX Spark (~$4k), AMD's Strix Halo / Ryzen AI Max+ 395 (~$2 to 3k), and the M5 Max 128GB (~$5k) head to head (Hardware Corner). Prompt processing favors CUDA hard. But token generation lands at a surprisingly...
LeCun's AMI Labs raised $1.03B at $3.5B pre-money — Europe's largest-ever seed round. Investors include NVIDIA, Jeff Bezos, Samsung, Toyota. LeCun's thesis: LLMs are fundamentally wrong for intelligence because they learn from text, not the physical world. AMI will build "worl...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.