Skills
Switching from API to chat interface degrades a model as much as downgrading a full generation
An audit of ChatGPT, Claude and Gemini across seven systems and nine benchmarks covering general capability, social bias and sycophancy found API evaluations score 3.4 percentage points higher in accuracy and 2.1 points higher in test-retest agreement than the same benchmark run through the deployed interface. For ChatGPT the API-versus-interface gap exceeded the API-only gap between two adjacent model generations. Varying system prompts, sampling parameters and reasoning settings shifted behavior but did not reliably close the gap, which means every published benchmark number you use to pick a model is measuring a surface most of your users never touch.
↳ Follow the thread