Fetching from the wire…
Public story · 2026-03-23 · source-backed
Cloudflare added Moonshot AI's Kimi K2.5 to Workers AI on March 19, making it the first frontier-scale open-source model available on edge compute with a full 256K context window, multi-turn tool calling, vision inputs, and structured outputs (Cloudflare Blog). Cloudflare reports 77% cost savings ($2.4M/year on a single workload) versus mid-tier proprietary models, with prefix caching and session affinity exposed as first-class features.
Kimi K2.5 is a 1T-parameter MoE model (32B active, released January 2026) with agent swarm capability that can self-direct up to 100 sub-agents. The validation from Cursor is the real signal here — a screenshot circulating on r/LocalLLaMA (232 upvotes) shows Cursor officially designating Kimi K2.5 as the best open-source model available in its platform. A major IDE endorsing a Chinese lab's model over Meta and Mistral offerings is a meaningful shift.
The edge deployment angle matters for anyone building agentic workloads. Running a capable model on globally distributed infrastructure without managing GPU clusters changes the economics of agent deployment from "provision a GPU fleet" to "call a function." Cline CLI 2.0 already includes it as a free default. The model is also accessible on Labla AI and OpenRouter.
This is what "open model wins" looks like in practice: a Chinese lab ships a frontier model, an American CDN distributes it globally at a fraction of proprietary cost, and an IDE company endorses it over models from companies that raised ten times more capital.
Each link below shares sources, entities, or timing with this story.
Cursor shipped Composer 2 on March 19, marketing it as a proprietary in-house model. Within 24 hours, a developer found the API routing to kimi-k2p5-rl-0317-s515-fast. The model powering the most-hyped coding tool update of the month was Moonshot AI's Kimi K2.5 with continued...
The assumption that proprietary models own the coding benchmark crown just broke. Moonshot AI's Kimi K2.6 leads on 5 of 8 major agentic coding benchmarks while being the only open-weight model in the top tier. SWE-Bench Pro: 58.6% vs GPT-5.4's 57.7% and Claude Opus 4.6's 53.4%...
Raw coding capability is no longer a moat. Cursor just proved it with numbers that are hard to argue with. Cursor released Composer 2.5 on May 18, built on Moonshot AI's open-weight Kimi K2.5, a 1-trillion-parameter mixture-of-experts model that activates just 32 billion param...
Cloudflare announced Workers AI now supports large-scale model inference, starting with Kimi K2.5 (256K context, vision, multi-turn tool calling). Internal testing: switching a 7-billion-token/day security code review agent from a proprietary mid-tier model to Kimi K2.5 achiev...
The models are good. The license is the real story. Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. Th...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.