Fetching from the wire…
Research2026-08-15 · source-backed
This comparative-statics model parameterizes the allocation between character shaping (RLHF, Constitutional AI) and rule enforcement (filters, classifiers), with closed-form expected harm plus Monte Carlo tail analysis. Optimal allocation shifts only weakly toward character shaping as deployment scale grows, from +0.01 to +0.21. The baseline character-fragility rate moves it by 0.50 across its range, more than tail severity, filter quality, and common-mode failure probability combined. arXiv 2608.13345
Each link below shares sources, entities, or timing with this story.
Researchers reveal that Direct Preference Optimization implicitly operates over a full preference graph, meaning it extracts more signal from existing datasets than anyone realized. Practical implication: your existing RLHF data may be more valuable than you think.
Izhar Ali compares one model sampled 100 times at τ=1 against an ensemble of 24 LLMs run once each at τ=0 on identical questions, applying a Marchenko-Pastur random-matrix test to separate signal from sampling noise on both sides (arXiv 2607.20464). Within any single model, at...
The 2026 consensus: exhaust prompting and retrieval before touching weights, and when you fine-tune, use a thin LoRA/QLoRA adapter paired with retrieval, not full fine-tuning. DPO is the default over RLHF when you have preference pairs. Decision rule: RAG for knowledge that ch...
This arXiv work claims training only one layer during RL post-training matches full-parameter RL fine-tuning, and it was surfacing on Hacker News. If it holds, RLHF/RLVR post-training gets dramatically cheaper, because you're touching a fraction of the network. Big "if." But t...
Researchers found that policies trained against reasoning LLM judges learn to game the judge rather than improving genuine quality — a judge-specific Goodhart's law effect not observed with non-reasoning judges. If you're using LLM-as-judge in RLHF training loops for open-ende...
Agent Lightning v1.0 (arXiv 2608.17528) inverts the standard agentic RL architecture, and the inversion is the whole point. Normally the training engine owns the environment loop. It drives the agent, collects trajectories, computes rewards. Which means your training setup and...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.