SourcesOpus 4.6 SWE-bench Ceiling Analysissimonwillison.net·high signalXBlueskyLinkedInCopy linkAnalysis of Claude Opus 4.6 performance ceiling at 80.8% SWE-bench and what it means for the next generation of coding agents.SourceSource pagesimonwillison.net↳ Follow the threadStack layer / Threat patternBoris Cherny: production code written by Claude should clear a higher bar than human code, and he lists the guardrails Anthropic runsSimon Willison's WeblogStack layer / Update threadOpenRouter's automatic fallback silently changes model behaviour, and provider.only is the fixSimon Willison's Blog (citing Mohamed Moustafa)Stack layer / Threat patternClaude Code 2.1.269 Ships a Plugin Eval Runner and a Knob to Raise the Workflow Tool's Concurrent Agent Cap to 256Anthropic (claude-code CHANGELOG)Stack layerPython 3.15 soft-deprecates re.match() and adds re.prefixmatch() to say what it actually doesSimon Willison's BlogStack layer / Threat patternResearchers Say Coding-Agent Sandboxes Leak in Claude Code, Codex and Cursor, and Anthropic Took 50 Days to PatchUpstarts MediaStack layer / ContrastSakana AI Ships Fugu Max and Fugu Ultra v2, an Orchestrator That Routes Each Task to the Leanest Model, at $2/$6 per Million TokensSakana AIThreat pattern / Update threadDatasette ships security releases after an audit with Fable 5.1, GPT-5.6 and GPT-6 Astra found 'very subtle' permission bugsSimon Willison's WeblogPolicy dependency / Stack layerClaude Code's agent view adds a peek panel so you can answer a blocked session without leaving the listClaude Code Docs