ResearchAgentLAB First Benchmark for Long-Horizon Agent AttacksarXiv·high signalXBlueskyLinkedInCopy link5 novel attack types: intent hijacking, tool chaining, task injection, objective drifting, memory poisoning. 28 environments, 644 test cases.SourceSource pagearXiv↳ Follow the threadPolicy dependency / Stack layerMemSentry gates persistent memory writes on a signed security-state delta rather than on content classificationarXivPolicy dependency / Stack layerPattern: the agent config layer is being treated as an unmanaged dependency graph, and three independent sources said so this weekarXivPolicy dependency / Stack layerRetrieval that crosses into your dependencies' source, not just your repo, adds up to 6.3% pass@1 and survives version changesarXiv 2609.09987Stack layer / Threat patternAgentAudit attaches to a running agent and scores its trace on ten dimensions, exposing 95.1 vs 22.6 trust spreads at similar task completionarXivPolicy dependency / ContrastA Fine-Tuned 4B Qwen in 2.6 GB Beats GPT-5.6 on a Transit-Kiosk Agent Benchmark, and PEFT Gains Vanish by 27BarXiv 2609.10016Policy dependency / Stack layerCROSS-CATEGORY: Three Independent Agent-Action Gates Shipped in 48 Hours, All Judging the Command Against Stated IntentProduct Hunt, github.com/AGGIB/Stroq and rewarelabs.com (three independent sources; the 72% figure is Reware's own)Policy dependency / Threat patternAgents hit 80% F1 deciding whether a dependency CVE is exploitable, but fall below 70% explaining whyarXiv 2609.08040Stack layer / ContrastA weaker agent recovered 80% of a stronger proprietary agent's capability gap from black-box execution differences alonearXiv 2609.07131