Agents
Contamination-controlled study finds the coding harness barely matters, except on contest tasks where it swings 23.7 points
A private 256-task suite of repository and contest problems (arXiv 2609.11987, submitted 2026-09-08) tested whether vendor-native harness-model pairings beat a neutral harness. Aggregate differences were tiny: the Claude-native harness underperformed by 1.25 points on Claude Opus 4.8 (48.8% vs 50.0%) and the OpenAI harness overperformed by 1.25 points on GPT-5.5 (55.6% vs 54.4%). The split by task type is the real result: the native harness trailed by 9.0 points on repository tasks but led by 23.7 points on contest tasks, and the neutral harness cost 1.3-1.6x more per solved task on Opus 4.8. If your workload is repo-shaped, the harness you pay a premium for may be the wrong one.
Source
↳ Follow the thread