Emergence World ran 10 agents per world for 16 days and found no frontier model contained an injected attack — one acted on poisoned memory 46 hours later
IBM-affiliated researchers ran eight parallel persistent worlds of ten agents each from identical starting conditions (seven homogeneous frontier-model worlds, one mixed), generating 850,000+ LLM calls and nearly 50 billion tokens over 16 days before firing three stress events through ordinary interaction surfaces: indirect prompt injection, misinformation, and exposure of private agent memories. No world was resilient to all three, and detection did not equal containment — systems recognized threats while still writing adversarial content into persistent memory and acting on it up to 46 hours later. The mixed-population result is the sharp one for builders: the same model-persona pairing behaved substantially differently in mixed versus homogeneous populations, so model-level alignment is not compositional and per-model safety cards tell you little about the multi-agent system you actually ship.
Source
↳ Follow the thread