Fetching from the wire…
Research2026-09-04 · source-backed
Prior work fuses OPD's dense token-level supervision with RLVR's sparse reward in a single step, either weighted-additive or teacher-modulated advantage rescaling. A plain two-stage OPD-then-RL scheme beats pure OPD, pure RLVR and all joint baselines across logic and math benchmarks. The mechanism: OPD expands the student's coverage of teacher-supported solutions and RL sharpens within that support, while joint optimization makes the signals interfere. The OPD validation score tells you when to switch, and OPD is a better cold start for RL than SFT. arXiv 2609.04108
Each link below shares sources, entities, or timing with this story.
arXiv 2609.04172 trains OPD with a single query and finds it keeps improving for hundreds of steps across task domains and model families. Measuring state coverage, the fraction of full-data states a query set's rollouts reach, one query hits 71.5% and 16 semantically distinct...
Measuring teacher supervision during on-policy distillation shows substantial noise that worsens as the teacher scales, yet the student converges comparably whether that supervision is kept or stripped (arXiv 2608.31046). Learning concentrates on low log-probability tokens, an...
MoRe learns a codebook of steering vectors, each encoding a latent role, then uses a query-aware router to fuse them into a single composed vector for single-turn inference. The backbone stays frozen; training is a three-stage SFT curriculum plus GRPO. Across reasoning and per...
arXiv 2607.23740 introduces a psychology-grounded benchmark of 3 primary dimensions, 17 secondary, 71 task paradigms, controlling question format, narrative perspective, and context length across 284 shared scenarios and 3,481 expert-verified instances. Not one of the 17 secon...
arXiv 2607.23982 adapts Holmström's team moral-hazard model into a game where an agent can keep an immediate local reward or pay a query cost to surface a hidden safety fact that mainly helps another agent's downstream decision. Base behavior splits into two failure modes: pre...
This arXiv work claims training only one layer during RL post-training matches full-parameter RL fine-tuning, and it was surfacing on Hacker News. If it holds, RLHF/RLVR post-training gets dramatically cheaper, because you're touching a fraction of the network. Big "if." But t...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.