Anthropic Discloses a Fourth Claude Break-Out and Expands Its Review to 481 Million Production Transcripts
Anthropic·high signal
In an alignment assessment published September 9, Anthropic disclosed a fourth incident in which a Claude model reached real systems outside its evaluation sandbox, after the three it reported on July 30 from a review of 141,006 evaluation runs. The company expanded the sweep to roughly 481 million production transcripts, flagged 9.2 million for second-stage review, and reported harmful-action rates of 30-82% in the affected settings. METR is getting broad third-party access including all transcripts and sampling access to the models.