Fetching from the wire…
Public story · 2026-08-24 · high
Each section cites the paper, textbook, or codebase behind it, sourcing that keeps RL fine-tuning choices from becoming pure guesswork.
Why now: Published as of August 24, 2026, it's the single document I'd point someone to before they touch GRPO code, since it traces the sources instead of skipping to the algorithm.
Cameron Wolfe published a single document on reinforcement learning for large language models. It runs from REINFORCE to the methods behind frontier training, on his newsletter Deep Learning Focus. The throughline explains why GRPO-family training behaves the way it does, past the mechanics of the algorithm itself.
Wolfe credits the paper, textbook, or codebase behind each section, tying theory to code instead of leaving the lineage scattered across sources. RL fine-tuning hyperparameters often get set by guesswork because that lineage is hard to find in one place.
The post starts at first principles, then walks through the policy gradient algorithms behind current training runs. It moves into reasoning, agents, token efficiency, and reliability, with each section linking a longer writeup on that piece.
Wolfe names his sources directly: Nathan Lambert's RLHF Book, Sutton and Barto's textbook, and OpenAI's Spinning Up guide. He also cites Lilian Weng's notes on policy gradients and the TRL and OpenInstruct codebases.
Each link below shares sources, entities, or timing with this story.
40. HBR — AI intensifies work 41. Import AI #441 42. Simon Willison — AI hit piece 43. Cameron Wolfe — GRPO++ 44. CISPA — Moltbook study 45. arXiv — SkillRL 46. Google DeepMind — Gemini 3 Deep Think 47. OpenAI — Codex-Spark 48. Zhipu — GLM-5 ---
CDAO Cameron Stanley confirmed the DoD has begun engineering work on internally developed LLMs and expects them operational "very soon," with xAI and OpenAI as interim alternatives. Defense Secretary Hegseth's supply-chain-risk designation against Anthropic remains in effect....
Sumeet Vaidya (CEO of Crafting, previously Meta, Uber, Discord) argues the binding constraint on enterprise AI is resilience rather than speed, citing policy and access changes at Anthropic and OpenAI as instability teams underweight. His recommendations: design so you can swa...
She announced July 27 with July 29 as her last day, publishing the note she sent internally and citing seven months of health problems aggravated by startup pace, writing she's been unwell more than at any other point in her life. She joined Mira Murati's company in February 2...
An independent researcher ran the Political Compass across 16 models from Google, Anthropic, OpenAI, xAI, Meta, Mistral, Qwen and Kimi — 30 standard runs, 30 reverse-phrased, 10 reordered each, ~69,440 answers, with the scoring system reverse-engineered to control for ordering...
This is a supply-chain fact, and most people are still treating it as a geopolitics argument. Sequoia published "America's Open-Model Paradox" on July 24 with the number that reframes the whole conversation: Qwen's share of open-model fine-tunes went from 1% in January 2024 to...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.