Fetching from the wire…
Public story · 2026-08-31 · high
The XpertStation WS300 packs 748GB of coherent memory, but buyers on r/LocalLLaMA say the listed price gets you a contact form, not a cart.
Why now: The pricing and stock details surfaced in the r/LocalLLaMA thread as of August 31.
MSI listed its XpertStation WS300 at $99,999, built around Nvidia's GB300 Grace Blackwell Ultra desktop superchip. For a solo developer or small lab, 748GB of coherent memory on a desk beats renting cloud GPUs, if anyone could buy one.
The spec sheet is real. 252GB of HBM3e runs at 7.1 TB/s, backed by another 496GB of LPDDR5X at 396 GB/s. Dual 400GbE networking rounds it out, all in a workstation case instead of a rack.
Buying one is a different story. The r/LocalLLaMA thread tracking the listing found the Newegg page shows out of stock. MSI's own product page swaps the buy button for a "Get Pricing" contact form. One commenter said they'd gotten quotes at or above $100k when they asked. Another called the whole thing a paper launch.
Asus, Dell, Gigabyte, Supermicro and HP all have GB300 DGX Station orders open too, with shipping promised in the coming months. Five vendors selling the same silicon at once means MSI's $99,999 is one bid among several, and the others haven't posted numbers yet.
Each link below shares sources, entities, or timing with this story.
Jensen Huang disclosed 100% of NVIDIA uses Claude Code, calling it "the first agentic model." Vera CPU: 88 custom Olympus cores, 1.2 TB/s LPDDR5X, paired with Rubin GPUs at 1.8 TB/s coherent bandwidth. 22,500+ concurrent CPU environments per rack. Dell, HPE, Lenovo, Alibaba, B...
paddo.dev makes the most contrarian read: the letter's substance isn't openness but paragraph nine, defending distillation as "a widely used technique for model improvement" and urging policymakers against "conflating legitimate model development techniques with misappropriati...
NVIDIA released Star Elastic, a post-training method that nests three submodels (30B, 23B, 12B) inside a single Nemotron Nano v3 checkpoint. The technique uses only 160B tokens (360x reduction vs pretraining) and cuts memory for deploying all three from 126.1GB to 58.9GB in BF...
3B active parameters, beats Qwen3.5-35B-A3B on AIME 2025 (92.4 vs 91.9), LiveCodeBench v6 (87.2 vs 74.6), and surpasses the larger Nemotron-3-Super-120B. Available on Ollama and HuggingFace under open license. Source
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
This is a supply-chain fact, and most people are still treating it as a geopolitics argument. Sequoia published "America's Open-Model Paradox" on July 24 with the number that reframes the whole conversation: Qwen's share of open-model fine-tunes went from 1% in January 2024 to...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.