Fetching from the wire…
Public story · 2026-08-24 · high
Privacy scores barely tracked with safety or security, and one model's robustness collapsed from 56.9 to 2.6.
Why now: The paper went up on arXiv on August 20, giving builders comparing open-weight models a rubric beyond one safety number.
A testing platform called aiXamine ran more than 120 large language models through 46 safety, security and privacy checks.
It found that tighter alignment makes models refuse more often. One model that scored 99.3 on safety alignment turned down one in three benign queries during testing.
A model's privacy score barely moved with its safety or security score in the same testing. The three risks don't rise or fall together, so a high score on one doesn't predict the others.
The paper treats safety, security and privacy as connected, not separate rankings. It ran the checks across nine services, in more than 5,000 test runs total.
One model was distilled off-policy, without a correction step back onto the teacher's own outputs. Its robustness score dropped from 56.9 to 2.6 on the same base architecture.
A related quantization study found a 4-bit compressed model beating its uncompressed teacher on most benchmarks. Compression trade-offs cut in different directions depending on the method.
Each link below shares sources, entities, or timing with this story.
A blinded judge checks root cause and impact against 95 real CVEs, and no frontier model made the ten-model lineup.
It automates the data-flow, crash-semantics, and commit-history work engineers do by hand.
The errors trace back to how the benchmark pairs pull requests with GitHub issues, not just to model quality.
A training-free fix called ChannelGuard held steady across three model backends, filter or no filter, blocking every tool-poisoning attempt.
One model scoring 99.3 on safety alignment refused one in three benign queries. It also documents distillation-induced robustness collapse, where off-policy distillation without on-policy correction dropped robustness from 56.9 to 2.6 on the same base architecture. (arXiv 2608...
Llama caught 80% of borderline anomalous logins in testing, versus 20% for Wazuh and 15% for OpenSearch.
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.