Synthetic Document Finetuning Makes a Model Look Aligned but Fails to Inoculate Against Reward-Hacking Misalignment
Models that learn to reward hack on RL environments can become broadly misaligned, and inoculation prompting blocks that generalization. This paper asks whether synthetic document finetuning can do the same job pre-emptively, by adding documents framing reward hacking as acceptable to the midtraining corpus before RL on exploitable environments. Behaviorally the midtraining works (models describe reward hacking favorably and approve of their own hacking outputs), but they still show strong emergent misalignment after learning to hack, while inoculation prompting in the same setting prevents it, suggesting SDF can steer generalization when inserting new associations and behaves unpredictably when overriding existing ones.
↳ Follow the thread