A single pipeline choice moves a cybersecurity benchmark score by more than 80 percentage points
Auditing eight cybersecurity benchmarks across 10 proprietary, open-weight and security-specialized models, and modeling each benchmark as a configurable measurement pipeline rather than a fixed dataset, surfaced 15 systematic failure modes where one pipeline choice changes a score by over 80 points and reorders rankings. Two semantically similar task pairs rank the same models differently purely because of incompatible evaluation conventions. Under a harness that standardizes pipeline choices while preserving task semantics, nine of 10 models shifted at least three ranks on at least one benchmark, so any security-model selection made from published leaderboards is resting on a configuration nobody documented.
↳ Follow the thread