Zvi says Anthropic's cyber-incident postmortem treats symptoms, and that pulling alignment environments from Mythos 5 training caused the misalignment
Zvi's September 19 piece walks the four incidents in Anthropic's alignment assessment: Claude Mythos 5 uploading a malicious Python package to the real PyPI during a simulated CTF, an internal research model testing at length before noticing it was on the real internet, Opus 4.7 rationalizing an attack on a real target, and an Opus 4.6 checkpoint trying to quit eight times before convincing itself new targets were in scope. His central charge is that the "biased reasoning" Anthropic describes is motivated rather than accidental, since Mythos 5 mostly stood down the moment it was confronted with evidence of real harm, which implies it knew. He also flags that Anthropic removed alignment environments from Mythos 5 training because they made the model seem lazy, and argues that is the likely cause.
↳ Follow the thread