Policy dependency / Stack layer
Pattern: four coding-agent CLIs shipped sandbox or trust work in the same week
GitHub
Stack layer / Threat pattern
Abnormal AI runs Bedrock AgentCore Code Interpreter as an ephemeral scratch pad at billion-message scale
AWS Machine Learning Blog
Stack layer / Threat pattern
Agent Frameworks Detect Dangerous Plan Steps and Then Execute Them Anyway; Fewer Than 20 Lines Closes the Gap
arXiv 2609.15293
Policy dependency / Stack layer
A Bun packaging quirk left a vulnerable undici in Cline after the CVE was supposedly remediated
GitHub
Stack layer / Contrast
Tip: gap-trap's "Proven Red" gate runs every new test against the old code and fails when it passes
GitHub
Stack layer / Update thread
Splitting a CTF Task Into Isolated Sub-Contexts Lets a Local gemma-4 Solve 18.52% of Challenges Standard Agent Loops Fail
arXiv 2609.12839
Stack layer / Follow-up thread
Ericsson's Multi-Agent Code Reviewer Hits 96% Accuracy, With a Third of Its Correct Findings Rated Must-Fix
arXiv 2609.15877
Stack layer / Contrast
Claude Code 2.1.271 Adds `claude plugin eval`, a Scored Reproducible Test Suite for Your Own Plugins
Releasebot (Anthropic Claude Code changelog mirror)