ResearchSkillCraft tool composition benchmarkWweb·medium signalXBlueskyLinkedInCopy linkBenchmark for agent skill abstraction and cross-task reuse. Skill saving/reuse reduces token usage by 80%. Success rate correlates with tool composition ability. Code released. arXiv 2603.00718↳ Follow the threadPolicy dependency / Stack layerA Bun packaging quirk left a vulnerable undici in Cline after the CVE was supposedly remediatedGitHubPolicy dependency / Stack layerRIPPLE: an edit confined to one prompt-policy segment changes downstream behavior, so replay candidate edits after previously accepted ones before persistingarXiv 2609.12127Policy dependency / Stack layerPraisonAI's MCP auth policy validated credentials for api-key and bearer but waved through basic and OAuthNVDStack layer / Threat patternSnyk put its agent-skill scanner behind a free web page called Skill InspectorSnyk LabsStack layer / Threat patternA 50-PR code review benchmark puts GPT-5.6 Luna at 28x cheaper than Astra and 74% precision against 96%EntelligencePolicy dependency / Stack layerCodeBLEU Scored 91% for Both RAG Strategies While One of Them Hallucinated APIs 56.4% of the TimearXiv 2609.12464Stack layer / Threat patternSkillSecurer Finds Latent Prompt-Injection Vulnerabilities in More Than 17% of Popular Published Agent SkillsarXiv 2609.14079Policy dependency / Stack layerHazardAuditor runs Claude Code, Codex, Hermes and OpenClaw in one harness and normalizes their events to train a guard modelarXiv / HuggingFace Daily Papers