Fetching from the wire…
Infra2026-06-28 · source-backed
June local-inference benchmarks across the 128GB class put NVIDIA's DGX Spark ($4k), AMD's Strix Halo / Ryzen AI Max+ 395 ($2 to 3k), and the M5 Max 128GB (~$5k) head to head (Hardware Corner). Prompt processing favors CUDA hard. But token generation lands at a surprisingly close 34 to 38 tok/s on 120B models. Strix Halo's load-bearing number is ~180 GB/s real usable bandwidth, and the community runtime is Vulkan via llama.cpp, not ROCm. Translation: prompt-heavy agentic loops want a CUDA box, but chat-style generation on the cheaper AMD option is genuinely competitive. Buy for your actual workload shape, not the headline number.
Each link below shares sources, entities, or timing with this story.
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Build b10677 fixes ggml_vk_graph_optimize, where is_src_of didn't treat two views of one tensor as dependent, so the optimizer reordered nodes across aliased reads and writes. Maintainers describe the result as silently wrong tokens under greedy decoding, different output on e...
Reuters, via Tech Startups, reports capital released against deployment milestones with Anthropic deploying up to two gigawatts of Instinct MI450 starting 2027. Same structure as Nvidia/OpenAI: compute vendor capital flowing to the lab that commits to buy the silicon. A two-gi...
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
The French startup argues the belief that GPUs suit agentic workflows poorly is a misconception, and says its Kog Inference Engine reaches up to 30x decoding speedups on standard NVIDIA and AMD datacenter GPUs with no new hardware. The approach is hardware-aware optimization o...
The end-of-summer update covers CUDA, ARM64, Metal and Vulkan backends for all core engines, experimental engines for music and 3D asset generation, and a router doing semantic and policy routing. It installs as a single OS service managing models and engines behind one base U...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.