Fetching from the wire…
Public story · 2026-08-26 · high
Ox Alpha, the OpenRouter model that more than doubled DeepSeek's usage before Z.ai claimed it, ships open weights tonight.
Why now: Z.ai confirmed authorship to Bloomberg on August 26, ending a guessing game that had run since the weekend.
Ox Alpha showed up on OpenRouter over the weekend with no listed creator. Usage more than doubled DeepSeek's within days, making it the platform's biggest launch with no name attached.
The mystery drew a crowd. A DeepMind researcher guessed Gemini. A browser-based fingerprinting tool pointed at the GLM family at six of nine probes.
On August 26, Z.ai told Bloomberg it built Ox Alpha as a new GLM-series iteration, with open weights going out tonight. On OpenRouter it's listed as stealth/ox-alpha, a reasoning model for coding and agentic work with a 1,048,576-token context window. Pricing for prompts and completions is zero, and it takes text, image and video input.
The number I trust here comes from a repo test, not a leaderboard score. Cline ran Ox Alpha and Fable on the same bug from its own codebase, and both fixed it. Ox Alpha used far fewer thinking tokens to get there.
Cline's test thread noted that Fable announced it had found the root cause seven separate times before it acted on the fix. That's a measured cost difference from real use.
Don't take fewer reasoning loops as proof of better reasoning. A thread on r/singularity raised the fair objection that committing to a first conclusion faster can also mean skipping edge cases.
The "80% coding score" that circulated last week dropped to 58.4% once someone ran the full evaluation. Any single number tied to this model deserves a second look once the weights are public.
If the weights ship tonight as promised, this becomes the cheapest way to test long-context agentic workflows. The free OpenRouter listing and million-token context remove the cost excuse most people have used to skip that test. Pin the exact model string before you test, since free stealth listings on OpenRouter don't usually stay free.
Each link below shares sources, entities, or timing with this story.
I've spent the last year assuming that if I wanted real agentic coding quality, I paid for a closed model. That assumption took a hit on June 1. MiniMax shipped M3 with a new sparse-attention architecture (they call it MSA) that handles up to 1M tokens at roughly 9x prefill an...
OpenRouter listed stealth/ox-alpha on August 20 with a 1M context window, free pricing, and a single anonymous third-party provider, described as a reasoning model for coding, sustained agentic work, and text-plus-visual production workloads. OpenRouter is explicit that it isn...
Zhong, Raghunathan, Laidlaw and Steinhardt fed 280 identities through Claude Code across four tasks. Against recognized safety researchers versus general users, Claude dropped behavioral confidence 1.4pp, increased reasoning usage 4.0pp and graded 0.11 points harder. Being tol...
Weeks after launch, Z.ai's open-weight GLM-5.2 now accounts for roughly 75% of all Z.ai model traffic on OpenRouter, with at least one provider serving it past 125 tokens per second (GIGAZINE, citing OpenRouter). The numbers behind the surge: an Artificial Analysis Intelligenc...
modelprint runs 9 infrastructure probes against any OpenAI-compatible endpoint from a static page with no server, keys never leaving the tab. Its day-one run against 12 candidates scored stealth/ox-alpha at 6 of 9 probes and 4 of 4 normalized tokenizer counts matching z-ai/glm...
It showed up August 20 with 1,048,576 tokens of context, 131,072 max output, and text, image and video input. Stripe's Patrick Collison called it "very impressive," no lab has claimed it, and speculation splits between Z.ai's GLM family, Microsoft's MAI and Xiaomi's MiMo. The...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.