Human Annotators Penalize an Agent's Extra Conversation Turns Twice as Hard as Extra Tool Calls
arXiv 2609.17985 (16 Sep 2026) argues task success alone hides agents that frustrate users through repeated questions, redundant searches and avoidable revisions. RideWay is a stateful tool-calling ridehailing benchmark paired with Efficiency Utility, a success-gated metric discounting trajectories for excess tool calls and user-facing turns against task-specific reference effort, with penalties calibrated by human paired preferences. Across 58 tasks and 24 models the fitted penalty for excess turns is about twice that for excess tool calls, and on held-out preferences Efficiency Utility hits 78.7% accuracy overall and 90.6% when trajectories differ in turns, but chance level when they differ only in tool calls, which marks the limit of count-based tool-use evaluation.
↳ Follow the thread