Fetching from the wire…
Top 5 · 2026-07-16 · source-backed
Mira Murati's lab finally shipped a full LLM, and it's Apache 2.0.
Inkling is 975B total parameters with 41B active in a MoE configuration, multimodal on input (text, image, audio) and text out, trained on 45 trillion tokens. The context number is the fun part: 1M tokens in the open weights, 256K via their own API and Tinker. The open release is more capable than the hosted one on context length, which is a choice I'd love to hear the reasoning behind. A smaller Inkling-Small (276B-A12B) ships alongside. Weights are on Hugging Face, with day-one availability on Databricks, Baseten, Modal, and vLLM. Hugging Face has the launch post.
The architecture is genuinely odd in ways worth reading the card for: non-RoPE relative positional encoding, a 5:1 local-to-global sliding-window attention ratio, short convolutions, and 2 shared expert sinks. That's a lot of deviation from the standard recipe for a first release. TML clearly spent their independence on architecture research rather than scaling the known thing, which tracks with everything the lab has said about itself.
Now the part that matters strategically. Inkling ranks #41 on the Intelligence Index. That makes it the leading American open-weights model. It also puts it behind GLM-5.2 and Kimi K2.6. The best open model America has is third, behind two Chinese labs. It does hit #9 in the Agentic Web App Arena with notably concise reasoning and strong tool calling, which for builders is arguably the more relevant number, since agentic tool use is what you're actually deploying.
This connects directly to the Siegel Endowment paper published in Fortune this week, arguing governments and nonprofits should fund open source AI. It hit 243 points and 86 comments on Hacker News, and Inkling is the argument's best exhibit. Every significant open-weight release to date has been a byproduct of some company's competitive strategy. Meta's, Alibaba's, Mistral's, now TML's. Strategy changes. When Meta decided open weights no longer served them, the open ecosystem lost its anchor overnight. Siegel's point is that an ecosystem whose existence depends on the strategic convenience of four companies isn't an ecosystem, it's a marketing budget.
What builders should do: pull Inkling-Small (276B-A12B) before you touch the big one. 41B active on the flagship still means you need serious hardware, and the small variant is where you'll find out whether the architecture's quirks help or hurt your workload. The 1M context in open weights is the real gift here, because that's the constraint that usually forces you back onto a hosted API. If you've got a document-heavy pipeline that's been paying per-token for long context, this is worth a weekend of benchmarking.
Also, hold the "American open weights are back" framing loosely. Third place is third place.
Each link below shares sources, entities, or timing with this story.
This is a supply-chain fact, and most people are still treating it as a geopolitics argument. Sequoia published "America's Open-Model Paradox" on July 24 with the number that reframes the whole conversation: Qwen's share of open-model fine-tunes went from 1% in January 2024 to...
Moonshot AI dropped Kimi K2.7-Code on Hugging Face on June 12. The specs are loud: 1T-parameter MoE with 32B active across 384 experts, a 256K context window, Modified MIT license, tuned for long-horizon agentic software engineering (MarkTechPost). Moonshot reports +21.8% on K...
The essay argues closed frontier labs risk concentrating too much power in too few hands, framing American open weights as the answer to Chinese open-source models (FT, corroborated by CNBC and Fortune). The sharpest lines: "I do not understand why anyone who believes that AI...
An open-weight Chinese frontier model is now a dropdown option in Microsoft's coding product. That happened before anyone finished characterizing what the model does. GitHub's changelog dated August 6 makes Kimi K3 generally available across Copilot Pro, Pro+, Max, Business an...
Xiaomi released MiMo-V2.5-Pro, a 1.02 trillion parameter mixture-of-experts model (42B active) with 1M token context, fully MIT licensed. In benchmarks, it achieves 63.8% success on agentic tasks using 40-60% fewer tokens than Claude Opus 4.6 or GPT-5.4 for comparable results....
The models are good. The license is the real story. Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. Th...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.