Sources
Ryan Lopopolo: you can only evaluate an agent in domains you already know, and the model was trained on what non-experts rewarded
This 2026-09-12 essay names a failure mode that shows up in every long agent run: outside your own expertise you have nothing to check the model against except its priors, and those priors were shaped by non-experts rewarding output that experts would call poor. His running example is defensive exception handling and other code patterns that read as competent to a reviewer who cannot tell, and which compound across iterative work products where nothing enforces long-term coherence. The Punnett-square framing is the sharp bit: an observer concludes the AI is competent whether or not they understand the domain, so perceived competence carries no information.
↳ Follow the thread