VoicesWillison 5 Posts: SWE-bench Analysis GGML HuggingFace Gemini 3.1 Prosimonwillison.net·high signalXBlueskyLinkedInCopy linkWillison covered SWE-bench Feb leaderboard (Opus 4.5 leads 76.8%), GGML joins HuggingFace, Gemini 3.1 Pro, Taalas 17K tok/s custom silicon. Called GGML joining HF hard to overstate.SourceSource pagesimonwillison.net↳ Follow the threadPolicy dependency / Stack layerPaul Christiano Joins the OpenAI Foundation Board and Its Safety and Security CommitteeOpenAI BlogPolicy dependency / Stack layerCROSS-CATEGORY: Three Independent Agent-Action Gates Shipped in 48 Hours, All Judging the Command Against Stated IntentProduct Hunt, github.com/AGGIB/Stroq and rewarelabs.com (three independent sources; the 72% figure is Reware's own)Stack layer / ContrastApple's A20 Pro widens the iPhone memory bus to 96-bit, putting on-device inference near 115 GB/sr/LocalLLaMA (corroborated by Notebookcheck and 9to5Mac)Threat pattern / Update threadDatasette ships security releases after an audit with Fable 5.1, GPT-5.6 and GPT-6 Astra found 'very subtle' permission bugsSimon Willison's WeblogStack layer / Update threadSpotify shipped a Claude Code plugin claiming 90% token savings, and r/ClaudeAI concluded it reinvented subagentsr/ClaudeAI (Spotify Engineering blog, 2026-09-03)Stack layerKernel fusion took GLM 5.3 Flash from 29 to 40 tok/s on an M3 Ultra by removing idle GPU gaps, not by better matmulsr/LocalLLaMA (github.com/IngeniousIdiocy/ds4)Stack layer / ContrastIBM released Granite Time Series PatchTST-FM-r2 with a commercial-friendly licenseHugging Face Blog (IBM Research)Policy dependency / Stack layerAnthropic Gave EU Cybersecurity Agency ENISA Access to Mythos 5, Three Months After Release, and Still Withholds 5.1Bloomberg