Fetching from the wire…
Public story · 2026-08-04 · high
Antares-3B, built on IBM Granite, approaches GPT-5.5 on vulnerability localization at under $0.002 per task.
Why now: The paper posted to arXiv on August 3, giving agentic vulnerability scanning a concrete per-task price for the first time at this size.
Antares-3B closes in on GPT-5.5's vulnerability-localization accuracy, using a 3-billion-parameter model built on IBM's Granite base, per the arXiv paper posted August 3.
That's the story for builders: a full 500-task evaluation sweep finishes in about 15 minutes on one H100. Amortized, that's under 2 seconds and under $0.002 per task, cheap enough for continuous scanning inside a CI budget, not a frontier-API one.
Training combines supervised fine-tuning on cybersecurity reasoning and repository-exploration data with reinforcement learning from verifiable rewards run over vulnerable repositories, per the paper.
It ships in three sizes: 350M, 1B, and 3B parameters.
Size-adjusted, the 3B version beats open-weight models 200 times larger, per the paper.
The paper doesn't say whether the gains hold on vulnerability classes outside its training repositories or on codebases in unfamiliar languages.
A related paper found that post-training on office workflows, with no software-engineering tasks included, still improved SWE-Bench Pro results. It's more evidence that domain-specific training data can matter more than raw parameter count.
At that price, running vulnerability localization on every pull request costs less than the CI minutes it takes to execute. That turns it from a paid security vendor's product into a default CI step, assuming the accuracy holds outside the paper's benchmark repos.
Posted to arXiv August 3, the paper puts a per-task price on agentic vulnerability scanning at this model size for the first time.
Each link below shares sources, entities, or timing with this story.
A blinded judge checks root cause and impact against 95 real CVEs, and no frontier model made the ten-model lineup.
Built on IBM Granite bases with supervised fine-tuning on cybersecurity reasoning plus repository-exploration data, then RL from verifiable rewards over vulnerable repos. Antares-3B outperforms open-weight models over 200x larger. The economics are the story: a full 500-task e...
Privacy scores barely tracked with safety or security, and one model's robustness collapsed from 56.9 to 2.6.
It automates the data-flow, crash-semantics, and commit-history work engineers do by hand.
Llama caught 80% of borderline anomalous logins in testing, versus 20% for Wazuh and 15% for OpenSearch.
An ablation that skipped the router entirely tied the full system's score, per the paper.
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.