Fetching from the wire…
Public story · 2026-07-01 · high
DeepReinforce's MIT-licensed Ornith-1.0 fits 35B parameters into a 20GB file that runs on one high-memory machine.
Why now: Willison published his write-up on June 29, giving developers a dated, working example of a self-training local coding model to test against hosted tools.
Simon Willison ran Ornith-1.0, an open coding model that writes its own reinforcement-learning training scaffold, on a single machine, per his June 29 write-up.
That's the case DeepReinforce is making to developers wary of sending code to a hosted API. The full agent loop, tool calls included, runs from a 35B mixture-of-experts model packed into a roughly 20GB file, on one high-memory machine you own.
Willison loaded the GGUF file in LM Studio and wired it into his own Pi coding harness. In his tests, Ornith-1.0 drove an agent loop across many tool calls without stumbling, he wrote. The model builds on two pretrained bases, Gemma 4 and Qwen 3.5, and ships under an MIT license.
The self-scaffolding part is the novelty: the model writes the training harness that guides its own reinforcement-learning improvement. It's a credible local fallback for agentic coding, especially for anyone who read the steganography story and wants an alternative to hosted models.
Each link below shares sources, entities, or timing with this story.
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Vicki Boykis wrote a post titled exactly that, "Running local models is good now," and it hit 1,437 points on Hacker News with 551 comments. Her claim is specific and checkable. Gemma 4, the gemma-4-26b-a4b and gemma-4-12b-qat variants, runs agentic coding at roughly 75% of fr...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
Hugging Face published its Summer 2026 State of Open Models report on August 14, and one statistic in it went almost entirely unremarked in the coverage. By July 2026, agents rather than humans became the Hub's primary users. Claude Code alone accounted for 44.4% of all agent...
Xiaomi released MiMo-V2.5-Pro, a 1.02 trillion parameter mixture-of-experts model (42B active) with 1M token context, fully MIT licensed. In benchmarks, it achieves 63.8% success on agentic tasks using 40-60% fewer tokens than Claude Opus 4.6 or GPT-5.4 for comparable results....
Google dropped Gemma 4 and it's not incremental. The 31B dense model ranks #3 on Arena AI with an ELO of 1,452, scores 85.2% on MMLU Pro, 89.2% on AIME 2026, and 80.0% on LiveCodeBench v6. It outperforms models 20x its size. Under Apache 2.0. At $0.20 per run. Only Opus 4.6 an...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.