Fetching from the wire…
Public story · 2026-08-05 · high
Its published coding scores were run across four different harnesses, so none of them can be compared to each other yet.
Why now: The cluster run surfaced in r/LocalLLaMA discussion covered in the August 5 briefing, while K3's benchmark table draws scrutiny for mixing harnesses.
A LocalLLaMA poster ran Moonshot's full Kimi K3 on 16 Nvidia GB10 chips and clocked 20-plus tokens a second, per a thread with 1,330 upvotes.
That's roughly $64,000 in hardware producing frontier-adjacent coding output at your desk. Dspark speculative decoding pushed bursts to 38 tokens a second, with prefill hitting 750. For anyone weighing a local rig against a Claude Code or Codex subscription, that's a real data point instead of a marketing slide.
Moonshot's scores land in frontier-adjacent territory: 88.3 on Terminal-Bench 2.1, 81.2 on FrontierSWE, 77.8 on ProgramBench raw pass, 67.5 on DeepSWE, 42.0 on SWE Marathon.
But those five numbers weren't run in one harness. Moonshot's own reporting mixes results across Kimi Code, Claude Code, Codex and mini-SWE-agent without saying which score came from which one. K3 is also reportedly sensitive to whether its thinking history gets preserved between turns, another variable the published numbers don't control for.
Anyone comparing K3's scores to Claude Code or Codex output should ask which harness produced them first.
Each link below shares sources, entities, or timing with this story.
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
Three moves, two days, no coordination between them. August 10–11: GitHub shipped Ollama as a BYOK provider inside Copilot for JetBrains (GitHub Changelog). Unsloth released Unsloth Desktop with a command literally named unsloth start claude, which points Claude Code and Codex...
Moonshot AI dropped Kimi K2.7-Code on Hugging Face on June 12. The specs are loud: 1T-parameter MoE with 32B active across 384 experts, a 256K context window, Modified MIT license, tuned for long-horizon agentic software engineering (MarkTechPost). Moonshot reports +21.8% on K...
Z.ai's GLM-5 (744B/40B MoE, MIT license, 205K context) is free on NVIDIA NIM at 40 req/min with no credit card. Benchmarks: 77.8% SWE-bench Verified (highest open-source), 56.2 Terminal-Bench 2.0 (approaching Opus 4.5's 59.3). Trained entirely on 100,000 Huawei Ascend chips. Y...
Terminal-Bench 2.1 results (entries dated June 17) put Codex CLI on GPT-5.5 first at 83.4%, Claude Code on Fable 5 second at 83.1%, and Claude Code on Opus 4.8 at 78.9%. The asterisk matters more than the ranking: Fable 5 and Mythos 5 have been export-suspended since June 12,...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.