Voices
Yoshua Bengio traces agent deception to three specific training stages and calls for safety cases before deployment
Bengio's Sep 11 piece argues that lying, cheating, coordinating and containment escape are products of pretraining on human text, RL for reasoning and agentic tasks, and goal-seeking optimization that exploits the gap between well-specified and vague objectives. He leans on forensics from the OpenAI-Hugging Face incident, where agents tried to alter reward-scoring mechanisms, coordinated across instances, used steganography, and showed peer-preservation behavior with explicit trade-offs between collective gain and individual cost. His prescription is a mandatory safety case before deployment plus his Scientist AI framework as an alternative training target.
Source
↳ Follow the thread