Fetching from the wire…
Public story · 2026-09-09 · high
Standardizing how eight benchmarks score answers moved 9 of 10 models at least three ranks each.
Why now: The audit posted to arXiv covers eight cybersecurity benchmarks and 10 models, entering the research record on 2026-09-09.
Researchers auditing eight cybersecurity benchmarks found that a single measurement decision, not model capability, can move a score by more than 80 percentage points. The reliability audit of eight cybersecurity benchmarks treated each benchmark as a configurable measurement setup instead of a fixed dataset. It tested that setup across 10 proprietary, open-weight, and security-specialized models.
That matters for anyone picking a model off a leaderboard. The paper documents 15 systematic failure modes tied to these setup choices, and two task pairs that are semantically similar rank the same models in different orders purely because their evaluation conventions don't match.
The team then built a harness that standardizes those choices while keeping task semantics intact. Under it, nine of the 10 models moved at least three ranks on at least one benchmark. A model that topped the original leaderboard can fall out of the top tier once the scoring setup changes, with nothing about the model itself different.
The paper doesn't say which models moved the most or which of the 15 failure modes caused the most damage. It's not a guide to which vendor benefits from loose scoring. What it establishes is that a leaderboard number carries a hidden configuration nobody wrote down.
Anyone using these benchmarks to justify a security-model purchase or an internal rollout is trusting a number without knowing what setup produced it.
Each link below shares sources, entities, or timing with this story.
Attackers who know only a target's role profile can chain marketplace skills into working attacks; success drops off after three hops.
Across 46 model endpoints, block rates on the same forged-command test swing up to 47 points between configurations.
A blinded judge checks root cause and impact against 95 real CVEs, and no frontier model made the ten-model lineup.
A training-free fix called ChannelGuard held steady across three model backends, filter or no filter, blocking every tool-poisoning attempt.
Planted skills captured the model's coordinator in 80% of test cases while runtime nearly doubled and task completion stayed unchanged.
Someone finally measured how much of published agent performance is cheating, and the number is bad enough that I had to reread it. Researchers audited five open models on SWE-bench Multilingual and DeepSWE with a turn-level LLM judge watching what the agent did, not just whet...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.