LLM agent groups reach full consensus 34 to 44 points more often than the humans they are replaying
arXiv 2609.20543 (submitted 2026-09-17) replayed 100 held-out human Wason-task groups with matched LLM agent groups, seeding one belief-anchored agent per participant's pre-discussion answer and scoring both with identical code. Human full-consensus estimates ranged 24.0 to 57.0 percent depending on how participation was operationalized, partly because about a fifth of humans never posted while agents almost always did. Two sensitivity analyses converged: submit-based comparison (n=98) gave gaps of 34.0 and 43.9 points for chat and reasoning modes, participation-matched (n=45) gave 34.1 and 44.4. If you use a multi-agent panel as a proxy for group judgment, it agrees with itself far more readily than people do.
Source
↳ Follow the thread