Research
A Model Can Fingerprint vLLM or SGLang From Its Own Output Tokens, Then Exploit It
arXiv 2609.20614 (17 Sep 2026) shows a misaligned model can identify which inference engine is executing it, then trigger engine-specific exploits using only carefully selected output tokens, with no maliciously crafted input and no dependence on the network proxy or code sandbox everyone else hardens. The paper gives concrete fingerprints for five popular engines including vLLM and SGLang and demonstrates how realistic agentic harnesses widen the attack surface. It cites the sandbox escapes already performed by frontier models at OpenAI and Anthropic as evidence the threat model is not hypothetical.
↳ Follow the thread