Fetching from the wire…
Public story · 2026-08-31 · high
The recipe stacks four cost cuts on consumer RTX 5090s and gets close to Qwen2.5-1.5B without a data center.
Why now: Puro-2B's paper posted to arXiv in August 2026, making the training recipe public for the first time.
Puro-2B trains a 2B parameter language model from scratch on consumer RTX 5090 GPUs for under $6,900, per the Puro-2B paper on arXiv. It runs pretraining in FP8 across up to 1.4 trillion tokens and gets close to Qwen2.5-1.5B under the authors' own evaluation protocol.
The authors cite more than $1.5 million to train Llama-3.2-3B and more than $700,000 to reproduce SmolLM3-3B. Puro-2B's budget is more than 200 times below the Llama figure, on hardware a hobbyist can own instead of renting from a cloud provider.
No single trick explains the gap. The savings stack four techniques: RTX 5090s over data-center accelerators, FP8 precision, hyperball optimization, and curriculum model averaging. Cut any one and the budget likely creeps back up, though the paper doesn't say by how much each piece contributes on its own.
What's missing is downstream testing. Coming close to Qwen2.5-1.5B on the authors' chosen benchmarks isn't the same as matching it on benchmarks a skeptical reader would pick. The paper doesn't say how the model behaves on tasks outside that evaluation set.
Each link below shares sources, entities, or timing with this story.
AMD unveiled its first rack-scale system to directly contest Nvidia at the rack level, with engineering samples in H2 2026 and mass production targeted Q2 2027. Microsoft joins Meta, OpenAI and Oracle as customers; Meta plans 1 gigawatt of Helios racks by year-end against a lo...
Meta Compute will sell excess AI compute and model access against AWS, Azure, and Google Cloud, backed by up to roughly $145B in 2026 infrastructure spend. CoreWeave shares fell 14% on July 8 as the market read Meta as a credible neocloud threat. For anyone renting GPUs, anoth...
Hugging Face published its Summer 2026 State of Open Models report on August 14, and one statistic in it went almost entirely unremarked in the coverage. By July 2026, agents rather than humans became the Hub's primary users. Claude Code alone accounted for 44.4% of all agent...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Sturgeon County, largely natural-gas powered, 3,000 construction and 300 operational jobs (Bloomberg). The detail I can't stop thinking about: Meta plans to rent idle GPUs to outside customers "like empty airline seats." That's a surplus-compute rental model emerging among hyp...
Meta is testing hardware from Watney Robotics, Kinova and ABB on cable swaps, server power-cycling and hardware reseating, against 2026 capital spending of $130B to $145B. Staff told reporters the concern is a shift away from experienced troubleshooting technicians toward lowe...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.