A Simulation Harness Gives Social-Reasoning Evals Ground Truth by Construction
Fuse (arXiv 2609.17496, submitted 15 Sep 2026) addresses the problem that social reasoning in advice settings has no verifiable ground truth: a target agent with a hidden motive interacts with other agents including one representing the user, who then consults the evaluated assistant to infer that motive, so the correct answer is known by construction. Simulation faithfulness was checked with a human study of 24,000 annotations. Applied to 12 LLMs, the framework shows user mediation compounds the difficulty, models are systematically sensitive to biased user framing, they need more detail than humans to reach a correct prediction, and longer conversations do not reliably help despite the chance to ask clarifying questions. Fuse and a 21,000-example dataset are open-sourced.
↳ Follow the thread