A 3.8B model trained from scratch to 0.384 CORE for $998 on eight B200s
Hugo Vergnes (via Hacker News)·low signal
Hugo Vergnes trained a Llama-style 3.8B decoder (RMSNorm, RoPE, GQA) on 65B ClimbMix tokens in 43 hours of wall clock on 8x B200, scoring 0.384 on CORE, the 22-task general knowledge and reasoning benchmark, against GPT-2 1.5B at 0.2565 and nanochat d32 at roughly 0.310. It hit the Hacker News front page today at 94 points. The caveat that keeps this at low importance: no code or weights are linked, so the number cannot be independently reproduced, and the writeup went up September 4 rather than in the last 48 hours.