Fetching from the wire…
Public story · 2026-02-22 · source-backed
No new posts today. Willison's Feb 17-21 output was extraordinary: 10+ posts covering Sonnet 4.6, GGML/HuggingFace merger, SWE-bench analysis (Opus 4.5 leads at 76.8%, Chinese models dominate top 10), and the Karpathy "Claws" amplification. His Showboat ecosystem — Rodney, Chartroom, datasette-showboat — remains the most concrete agent-tooling stack any individual developer has built. Expect resumed output Monday aligned with Anthropic's Briefing. simonwillison.net
Each link below shares sources, entities, or timing with this story.
Willison's prolific output continues: he covered ggml.ai joining HuggingFace, Taalas 17K tok/s, SWE-bench Feb 2026 leaderboard (Opus 4.5 leads at 80.9%, 4 Chinese models in top 10), Paul Ford's NYT essay, Martin Fowler's "LLMs eating specialty skills" observation, and Thariq S...
Willison published 8 posts on February 17 — his most prolific single day in recent memory. Key outputs: (1) Claude Sonnet 4.6 review, noting "similar performance to November's Opus 4.5" at Sonnet pricing, with SVG benchmark tests noting Sonnet 4.6 "consistently added decorativ...
Willison's same-day review provides independent verification of Anthropic's claims. His SVG generation benchmark and pelican test offer reproducible evaluation methodology. Also notable: his Showboat ecosystem (Rodney + Chartroom + datasette-showboat) represents the most compl...
The system card reports browser-agent injection falling from 31.5% to 3.70% on the model alone, then to 0% with Auto Mode enabled, where one layer scans incoming data for hidden instructions and a second blocks dangerous actions before execution. Gray Swan's independent genera...
If you've used Claude Code for any serious session, you know the drill. Approve. Approve. Approve. Approve. You stop reading the prompts after the fifteenth one. That's the worst possible security outcome, way worse than a well-designed automated check. Anthropic launched auto...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.