Letting agents exchange non-binding "cheap talk" before acting measurably stabilizes their policies across repeated games
Across four open-weight 7-9B models playing repeated Prisoner's Dilemma, Snowdrift, Stag Hunt and Harmony in six framings each, action policies drifted in all four games, with prevalence depending heavily on model and context. Agent-generated non-binding pre-play communication was predominantly stabilizing, with five corrected reversals concentrated in social or team framings, and controlled interventions on Qwen isolated two separable output-level channels: reduced action uncertainty and less between-round drift in action probabilities. The mechanistic result is the notable part — in Prisoner's Dilemma the authors identified a history-balanced policy-content direction in late transformer layers, and projecting it out increased realized switching during closed-loop play, showing the trajectory is causally sensitive to that component rather than merely correlated with it.
Source
↳ Follow the thread