Research
Frontier Models Score 1% Exact-Match on ICD-10 Coding When the Task Is Following a Real Manual End to End
TAM curates human-validated tasks from ICD-10-CM clinical coding and US federal sentencing, where each task requires following an authoritative manual containing tens of thousands of rules and executing interdependent steps across different sections to reach one exact answer. Retrieval-augmented generation, ReAct-style prompting, and an agent-harness baseline on GPT-5 all land at 1% exact match on ICD-10-CM and 15.5% on sentencing. The gap between this and saturated multi-hop benchmarks is a caution for anyone deploying agents against long procedural rulebooks.
↳ Follow the thread