VoicesWillison 5 Posts: SWE-bench Analysis GGML HuggingFace Gemini 3.1 Prosimonwillison.net·high signalXBlueskyLinkedInCopy linkWillison covered SWE-bench Feb leaderboard (Opus 4.5 leads 76.8%), GGML joins HuggingFace, Gemini 3.1 Pro, Taalas 17K tok/s custom silicon. Called GGML joining HF hard to overstate.SourceSource pagesimonwillison.net↳ Follow the threadStack layerPaul Ford on the paradox of democratized coding: accessibility clarified who shouldn't be relying on AI toolssimonwillison.netStack layer / Threat patternPattern: Anthropic's own bar is that Claude-written production code gets reviewed harder than human-written codeSimon WillisonStack layer / Threat patternSnyk put its agent-skill scanner behind a free web page called Skill InspectorSnyk LabsStack layer / Update threadsmolbenchmark ranks sub-8GB models by tokens per joule and thermals on hardware you already ownyuvrajsingh-mist.github.io (via r/LocalLLaMA)Stack layer / ContrastReal-SWE Benchmarks Coding Agents on Licensed Private Production Codebases: Fable 5.1 Tops It at 38.8%, GPT-6 Astra 33.8%Specific Labs / Hacker News (248pts, 137 comments)Stack layer / ContrastSix Claude Opus 5 Agents Benchmarked on CAD: OpenSCAD Produced Five Silent Geometry Failures, CadQuery Broke Loudly InsteadModelRift / Hacker NewsPolicy dependency / Stack layerA replay of 68,266 real Claude Code requests says plain LRU beats the clever KV-cache policiesGitHubPolicy dependency / Stack layerRIPPLE: an edit confined to one prompt-policy segment changes downstream behavior, so replay candidate edits after previously accepted ones before persistingarXiv 2609.12127