Agents
Harness planning text beats shuffled filler by 7.17 points, and a one-cent verifier catches 61 percent of bad episodes
arXiv 2609.20474 (submitted 2026-09-17) isolates what a harness actually contributes by pairing prewritten task-specific plans against shuffled policy text matched word for word, across 265 matched cells on tau^2-bench Retail plus an Airline pilot. Real guidance content improves oracle-verified success by 7.17 percentage points (90% task-clustered bootstrap interval 1.15 to 13.36), concentrated in harder tasks. A read-only terminal verifier rejects 61 percent of oracle-invalid Retail episodes while withholding 17 percent of correct ones, for under one cent per episode, so which component to buy depends entirely on what a wrong acceptance costs you.
Source
↳ Follow the thread