DispatchAI Gamestore Models Score Under 30% of Human PerformanceImport AI·high signalXBlueskyLinkedInCopy linkGPT-5.2 Gemini-2.5-Pro Claude Opus 4.5 all scored below 10% of human performance on 100 AI-generated game benchmarkSourceSource pageImport AI↳ Follow the threadStack layer / ContrastEmergence World ran 10 agents per world for 16 days and found no frontier model contained an injected attack — one acted on poisoned memory 46 hours laterarXivStack layer / ContrastByteShape's Qwen3.8-27B quants hit 99.63% of BF16 at 3.84 bpw, and argue KL divergence is the wrong quant metricByteShape (via r/LocalLLaMA, 126 upvotes)Policy dependency / Stack layer'Do this as quickly as possible' repeatedly got a Claude session flagged by a corporate security directorr/ClaudeAIStack layer / ContrastTypeSafe AI ships Jev, a model that returns typed probabilistic values instead of text, at $0.042 per million input tokens and free outputTypeSafe AIStack layer / ContrastPre-registered ablation shows removing an LLM verifier stage from an offensive-security agent shifts median reported findings from 0 to 2 per runarXivStack layer / Threat patternMemRiskBench Scores Long-Horizon Agent Memory Risks Deterministically, With No LLM Judge on the Pass/Fail PatharXiv 2609.14976Stack layer / Threat patternAgent Frameworks Detect Dangerous Plan Steps and Then Execute Them Anyway; Fewer Than 20 Lines Closes the GaparXiv 2609.15293Policy dependency / Stack layerOrdewell uses a read-only coding agent as planner, then spawns per-task agent sessions gated by a marker-based verdict engineGitHub