ResearchZeroDayBench frontier LLMs vs zero-daysWweb·medium signalXBlueskyLinkedInCopy linkGPT-5.2, Claude Sonnet 4.5, Grok 4.1 tested against 22 novel critical vulnerabilities. Frontier LLMs cannot autonomously discover/patch zero-days. ICLR 2026 Agents in the Wild workshop. arXiv 2603.02297↳ Follow the threadPolicy dependency / Stack layerPython's Import Statement Is an Execution Boundary: 90% of Initialization-Activated Advisory Vulnerabilities Are High or CriticalarXiv 2609.14791Stack layer / Threat patternSnyk put its agent-skill scanner behind a free web page called Skill InspectorSnyk LabsPolicy dependency / Stack layerHazardAuditor runs Claude Code, Codex, Hermes and OpenClaw in one harness and normalizes their events to train a guard modelarXiv / HuggingFace Daily PapersPolicy dependency / Stack layerA Bun packaging quirk left a vulnerable undici in Cline after the CVE was supposedly remediatedGitHubPolicy dependency / Threat patternA Fine-Tuned RoBERTa-Large Permission Gate Matches Claude Haiku 4.5 at Deciding What an Agent May ToucharXiv 2609.15422Stack layer / ContrastClaude Code lowers the medium dynamic-workflow guideline from 15 agents to 10, and defaults Pro plans to smallClaude Code ChangelogStack layer / ContrastVercel's AI SDK harness layer now authenticates nine coding agents through their own subscriptions instead of API keysVercel ChangelogStack layer / ContrastSplitting a CTF Task Into Isolated Sub-Contexts Lets a Local gemma-4 Solve 18.52% of Challenges Standard Agent Loops FailarXiv 2609.12839