Fetching from the wire…
Public story · 2026-08-22 · high
Llama caught 80% of borderline anomalous logins in testing, versus 20% for Wazuh and 15% for OpenSearch.
Why now: The paper is up on arXiv as of August 22, 2026, with the full accuracy, latency, and false-negative breakdown for all three tested models.
Meta's Llama 3.1 8B hit 89.3% accuracy detecting authentication anomalies in a new test, beating Wazuh's 52.0% and OpenSearch's 49.3%, per the arXiv paper.
Security teams that lean on Wazuh or OpenSearch for login monitoring are trusting tools that missed most of the true anomalies in this test. Wazuh's false negative rate was 68.6%, OpenSearch's was 74.5%, against 11.8% for Llama.
The researchers built their own testbed: endpoint-specific login logs spanning normal, borderline, and outright anomalous behavior, scored against one shared severity ground truth. Llama's F1 score came in at 91.8%.
The real gap shows up on the hard cases. On borderline anomalous logins, the ones that don't trip an obvious rule, Llama caught 80% of them.
Wazuh caught 20% of those same borderline cases. OpenSearch caught 15%.
Qwen 2.5 7B, the third model tested, wasn't the most accurate of the three. It had the lowest inference latency and produced valid structured output 100% of the time. That matters if the output feeds an alerting system instead of a person's eyes.
Neither Wazuh's rule engine nor OpenSearch's statistical model was built to weigh context the way a language model does, and the borderline numbers show it. The paper doesn't test live production traffic. It also skips the log volumes and attack types Wazuh and OpenSearch were actually tuned for in the field. A follow-up test against real production traffic would settle whether the gap holds outside a controlled testbed.
Each link below shares sources, entities, or timing with this story.
A training-free fix called ChannelGuard held steady across three model backends, filter or no filter, blocking every tool-poisoning attempt.
The errors trace back to how the benchmark pairs pull requests with GitHub issues, not just to model quality.
A blinded judge checks root cause and impact against 95 real CVEs, and no frontier model made the ten-model lineup.
The same researchers warn prompt wording matters almost as much as the retrieval backend, undercutting an easy graph-retrieval fix.
Privacy scores barely tracked with safety or security, and one model's robustness collapsed from 56.9 to 2.6.
The two judges scoring these 14,560 attacks disagreed by more than 3x on how often DeepSeek's agent partially complied.
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.