SourcesInternational AI Safety Report: Models Exhibit Behavioral DeceptionInternational AI Safety Report·high signalXBlueskyLinkedInCopy linkBengio-led report confirms sandbagging in frontier models. US declines to back report. First gov-grade deceptive alignment confirmation.SourceSource pageInternational AI Safety Report↳ Follow the threadPolicy dependency / Stack layerAgents hallucinate tools that do not exist, a 675B model does it as often as a 7B one, and merging MCP servers adds new failure surfacesarXiv 2609.19425Stack layer / ContrastMAGS routes coding-agent output through Dafny and reports 100% success at producing verified programs on 220 tasksarXivStack layer / Threat patternA Model Can Fingerprint vLLM or SGLang From Its Own Output Tokens, Then Exploit ItarXiv 2609.20614Stack layer / ContrastPACT benchmarks whether enterprise assistants break their own system-prompt rules under pressure from a persistent userarXiv / HuggingFace Daily PapersStack layer / ContrastOpenJev Reverse-Engineers TypeSafe's Classifier Inference, Then Renames Itself SemIf Mid-LaunchOpenJev / SemIf / Hacker News (638pts, 270 comments)Policy dependency / Stack layerLangChain ships a first-party integration that deliberately does not wrap the vendor's SDKGitHubStack layer / Update threadClaude Code Projects rebuilt around a coordinator that directs parallel threads, with per-thread model and thinking-effort settingsThe RegisterStack layer / Threat patternPlugin4Shell: One SHA-Pinning Bug Gives Zero-Click RCE in Claude Code, Codex, Copilot and Gemini CLIHelp Net Security (corroborated by The Register)