Reddit
A current OpenAI capabilities researcher goes public: models are too situationally aware to evaluate honestly
Dan Selsam, an OpenAI capabilities researcher since 2022 who helped pioneer chain-of-thought optimization there, published a personal statement on AI risk through Daniel Kokotajlo because he has no X account. His argument is narrower than the usual doom framing: models are becoming situationally aware enough that alignment evaluations no longer tell us how they would behave unobserved, so future experiments will teach us almost nothing new and models will increasingly seem aligned when they are not. He explicitly says pacing the frontier more carefully does not address this, which puts him at odds with the remedy the lab CEOs have been proposing.
Source
↳ Follow the thread