VoicesChollet ARC-AGI-3 Developer Toolkit Interactive Agency BenchmarkARC Prize·high signalXBlueskyLinkedInCopy linkFirst interactive reasoning benchmark. 2000 FPS local environments. Current AI systems score less than 5%. Measures adaptive intelligence.SourceSource pageARC Prize↳ Follow the threadStack layerThe same model scored 62.7% and 99.9% on ARC-AGI-3 depending only on the harness wrapped around itARC PrizeStack layer / Threat patternAlcaTRAz Defends Jailbreaks With Character-Level Perturbation Rules and No Model Access, Beating Llama Guard on 73.4% of CombinationsarXiv 2609.03693Policy dependency / Stack layerCopilot CLI 1.0.83 stable lets a custom agent list several models and fall through them in orderGitHubStack layer / Threat patternCrewAI 1.15.19 downgrades its own telemetry from core count to a coarse machine-size bandGitHubStack layer / ContrastCROSS-CATEGORY: Three Unrelated Orgs Shipped Coding-Agent Cost Routing in the Same 48 HoursGitHub Blog, Spotify Engineering and CodeRabbitStack layer / Threat patternBrockman says Astra was OpenAI's first training run on more than 100,000 GPUsStratecheryStack layer / Threat patternValidating a security patch by re-running the crash PoC inflates agent solve rates 1.83xarXiv 2609.04075Stack layer / ContrastThe Best LLM Catches 47% of Expert-Identified Requirement Defects While False-Flagging 11%arXiv 2609.03230