Fetching from the wire…
Top 5 · 2026-06-10 · source-backed
Cohere launched North Mini Code on June 9 under Apache 2.0, its first developer-focused model. The shape is the pitch: 30B parameters, mixture-of-experts, only ~3B active, and it runs on a single H100. It scores 33.4 on the Artificial Analysis Coding Index, competes on SWE-Bench Verified, SWE-Bench Pro, and Terminal-Bench 2.0, and claims up to 2.8x higher throughput with ~30% lower inter-token latency than Devstral Small 2. Cohere says it beats open models up to 4x its size.
Why this matters more than another benchmark post: this is a genuinely deployable local coding agent. Apache 2.0 means no usage restrictions, single-H100 means you can self-host it without a datacenter, and the throughput numbers mean it's fast enough to sit in an actual agent loop rather than a toy demo. Put it next to the NVIDIA RTX Spark Superchip hitting availability with 128GB unified memory and 120B-param local inference, and the "you need frontier infra for agentic coding" assumption is eroding fast.
I haven't run North Mini Code yet, so treat the throughput claims as vendor numbers until you A/B them on your own repo. But the use case writes itself. The work where you don't want to pay $50/M to Fable or ship proprietary code to an API: tight inner-loop edits, test generation, the high-volume agentic grunt work. Run a local 30B for the cheap 80% and reserve frontier calls for the hard 20%. That two-tier split, local model for volume, frontier for the genuinely hard problems, is starting to feel like the default architecture rather than a cost hack. Pull the weights, wire it into your harness, and measure it against whatever you're currently overpaying for.
Each link below shares sources, entities, or timing with this story.
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
Martin Alderson's essay "The upcoming AI margin collapse, part 1: GLM 5.2" hit 675 points and 462 comments on Hacker News, and it's the rare HN chart-topper that's actually about spreadsheet math instead of vibes. The argument is simple. Z.ai's GLM 5.2 delivers frontier-adjace...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill. SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-...
The open-weight race just changed constraint. Moonshot AI suspended all new consumer subscriptions on July 20, roughly 48 hours after Kimi K3 launched, because request volume pushed its compute cluster to capacity. Remaining GPUs are reserved for existing paid subscribers. Tec...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.