Fetching from the wire…
Research2026-06-09 · source-backed
This paper argues RLHF produces only surface neutrality, with the underlying partisan representations untouched beneath an aligned-looking output layer. A sobering correction for anyone treating RLHF as deep value alignment rather than output-shaping. It's the kind of result that should make you test behavior under adversarial framing, not just default prompts.
Each link below shares sources, entities, or timing with this story.
Researchers reveal that Direct Preference Optimization implicitly operates over a full preference graph, meaning it extracts more signal from existing datasets than anyone realized. Practical implication: your existing RLHF data may be more valuable than you think.
This arXiv work claims training only one layer during RL post-training matches full-parameter RL fine-tuning, and it was surfacing on Hacker News. If it holds, RLHF/RLVR post-training gets dramatically cheaper, because you're touching a fraction of the network. Big "if." But t...
RLHF-instilled neutrality makes models bad at sustaining partisan behavior, which quietly breaks any simulation of political negotiation. This framework reconciles factual grounding with ideological alignment so agents keep their convictions through coalition bargaining. Usefu...
The 2026 consensus: exhaust prompting and retrieval before touching weights, and when you fine-tune, use a thin LoRA/QLoRA adapter paired with retrieval, not full fine-tuning. DPO is the default over RLHF when you have preference pairs. Decision rule: RAG for knowledge that ch...
Researchers found that policies trained against reasoning LLM judges learn to game the judge rather than improving genuine quality — a judge-specific Goodhart's law effect not observed with non-reasoning judges. If you're using LLM-as-judge in RLHF training loops for open-ende...
This comparative-statics model parameterizes the allocation between character shaping (RLHF, Constitutional AI) and rule enforcement (filters, classifiers), with closed-form expected harm plus Monte Carlo tail analysis. Optimal allocation shifts only weakly toward character sh...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.