Research
A ReAct Agent Averaging 77% Per Run Succeeds All Five Times on Only 53% of Tasks
On AppWorld with GPT-4.1, a ReAct agent's per-run pass rate averages 77% but it succeeds in all five runs of the same task only 53% of the time, a 24-point shortfall the authors name the consistency gap. Their framework adds a Consistency Analyzer that pinpoints where a trajectory is likely to flip across executions and a Guideline Generator that converts the diagnosis into targeted guidelines committed to episodic memory and injected into future runs on similar tasks. That raises the fraction of tasks succeeding in all five runs by 16 points on same-task evaluation and 13 points on similar-task generalization.
↳ Follow the thread