Fetching from the wire…
OSS2026-08-14 · source-backed
alex000kim/nanoRL (MIT) scales from CartPole on a laptop to async distributed training on GPU clusters with vLLM rollout workers, across 7 files. The author optimizes explicitly for readability and forkability, excluding Megatron-scale parallelism and multi-tenant scheduling. 5 stars and 7 commits, so it's a teaching artifact, not a production trainer, which is exactly what makes it worth reading.
Each link below shares sources, entities, or timing with this story.
New batching algorithms enable ~7x, up to 12x+, longer-context GRPO training with no accuracy or speed penalty versus optimized FA3 and chunked-loss setups (Unsloth Docs). Qwen3-8B GRPO reaches 110K context on one 80GB H100 via vLLM plus QLoRA. For solo builders doing reasonin...
AxisAgentic (855 stars) is a runtime emitting append-only traces that reconstruct exactly what the model observed at any point, with explicit runtime markers for rollback, context compaction and discard events: which is what makes state-faithful SFT export possible without lea...
Gemma 4 12B dropped June 3, and the spec sheet is the kind of thing I read twice to make sure I wasn't misreading it. 11.95 billion params, Apache 2.0, reads text, image, audio, and video. No separate vision encoder. No separate audio encoder. The model handles all of it nativ...
The Segment co-founder published "Small models have arrived" on August 26, and it took 703 points on Hacker News (calv.info). His measurement: a personalized-news task that cost about a dollar on Sonnet-class models now runs at about a dime. Ten times cheaper, doing the job we...
Nathan Lambert doesn't hand out "step change" lightly, so when his June 22 Interconnects essay called GLM-5.2 "the step change for open agents," I read it twice. His argument is sharper than the usual "strong open model" take. Static intelligence benchmarks stopped mattering m...
The file is 1.56TB. That's the first thing you notice about the moonshotai/Kimi-K3 Hugging Face repo that went live today. 2.8 trillion total parameters, 104B activated, 896 routed experts with 16 selected plus 2 shared per token, 93 layers split 69 Kimi Delta Attention and 24...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.