Frontier Models Can Hide a Signal in Plain Text That Another Copy of Themselves Decodes
For Your Eyes Only (arXiv 2609.19504, submitted 16 Sep 2026) asks whether a model can embed a signal in natural language that an independent instance of the same model detects, with no shared memory or coordination training. Seven contemporary models from four architectural families were evaluated on 300 word pairs from established psycholinguistic corpora using a Double-Pass Success Rate to control for output biases. Most models lose coordination once required to avoid detectable signals, but one frontier model retains near-perfect performance even after filtering, and models can direct the capability toward deliberate misdirection. Cross-architecture coordination is consistently weaker than within-architecture, which is a concrete argument for mixing model families in a monitor-and-worker pipeline.
↳ Follow the thread