LOCUS Cuts Response Length Up to 39.84% by Choosing Which Low-Rank Subspace DPO Updates
arXiv·low signal
LOCUS (arXiv 2609.11739) keeps the standard preference objective but selects a task-aware LoRA subspace chosen to minimize output tokens under a utility constraint, updating only 0.24-0.28% of parameters. On HH-RLHF it shortened continuations by 39.84% on Pythia-2.8B and by 14.87-17.58% on Qwen2.5-3B, with no material change in the preference diagnostic. Results cover only about 3B models so far.