Fetching from the wire…
Research2026-05-11 · source-backed
Researchers reveal that Direct Preference Optimization implicitly operates over a full preference graph, meaning it extracts more signal from existing datasets than anyone realized. Practical implication: your existing RLHF data may be more valuable than you think.
Each link below shares sources, entities, or timing with this story.
Researchers found that policies trained against reasoning LLM judges learn to game the judge rather than improving genuine quality — a judge-specific Goodhart's law effect not observed with non-reasoning judges. If you're using LLM-as-judge in RLHF training loops for open-ende...
ARC-AGI-2 Reasoning Race. Gemini 3.1 Pro hit 77.1% and Claude Opus 4.6 hit 68.8% — both more than doubling their predecessors' scores in a single generation. The entire field moved ~10 points over the prior two years; each model jumped 30-40 points in one release. The top thre...
This arXiv work claims training only one layer during RL post-training matches full-parameter RL fine-tuning, and it was surfacing on Hacker News. If it holds, RLHF/RLVR post-training gets dramatically cheaper, because you're touching a fraction of the network. Big "if." But t...
The 2026 consensus: exhaust prompting and retrieval before touching weights, and when you fine-tune, use a thin LoRA/QLoRA adapter paired with retrieval, not full fine-tuning. DPO is the default over RLHF when you have preference pairs. Decision rule: RAG for knowledge that ch...
Decoder-only transformer with RMSNorm and rotary embeddings, using note-level tokenization NOTE(pitch, delta_onset, duration, velocity) so one forward pass advances the music by a complete note rather than an event fragment. Trained on a few hundred thousand MIDI files, ~300M...
This comparative-statics model parameterizes the allocation between character shaping (RLHF, Constitutional AI) and rule enforcement (filters, classifiers), with closed-form expected harm plus Monte Carlo tail analysis. Optimal allocation shifts only weakly toward character sh...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.