SourcesAgentShield First Open Benchmark 6 Agent Security Products 537 TestsAgentShield·high signalXBlueskyLinkedInCopy linkFirst open benchmark testing 6 agent security tools across 537 cases. Tool abuse detection is weakest category. Market optimizing for wrong threat.SourceSource pageAgentShield↳ Follow the threadStack layer / Threat patternalphaXiv shipped OpenResearch, a Rust runner for parallel research agents across any modelGitHub TrendingStack layer / Threat patternEdge0 runs a 35B MoE on Apple Silicon in 2.9 GB of active memory by streaming experts off SSDGitHubStack layer / Threat patternPydantic AI 2.42.0 adds a first-class provider for GitHub Copilot's OpenAI-compatible APIGitHubStack layer / Threat patternuv 0.12.11 now verifies source archives against uv.lock hashes before running their build backendsGitHubPolicy dependency / Stack layerMCP Inspector 2.6.0 raised its declared hono floor rather than trusting its lockfile, and documented whyGitHubStack layer / Threat patternRAGFlow 0.27.2 rewrites its Agentic RAG retrieval framework and patches a starlette CVEGitHubStack layer / Threat patternblock/goose 1.50.0 deletes fast model routing and the managed model registry for local inferenceGitHubStack layer / Threat patternPattern: git worktrees became table stakes across three agent harnesses in four daysGitHub