Fetching from the wire…
Public story · 2026-08-25 · high
A gzip-based compression test tied the free ox-alpha model to Z.ai's GLM-5.3, and a coding benchmark score backs it up.
Why now: OpenRouter's ox-alpha slot is still live and anonymous, the same listing this fingerprint targets.
Someone on OpenRouter fed ox-alpha, an anonymous model sitting in a free preview slot, its own system prompt back as a user message. The model broke character and named itself GLM, built by Z.ai, per dejan.ai's compression-distance analysis. If that holds, OpenRouter's free tier is running a full production model behind an anonymous label, not a cheap stand-in.
A single broken-character response isn't proof on its own, so the analysis added Normalized Compression Distance, a technique that fingerprints a model by how well gzip compresses its output using another model's output as a reference. Similar models compress tightly against each other. Different ones don't. The method needs nothing beyond text a model already returns.
Run across 60 prompts against GPT-5.5, Claude Opus 5, two Gemini variants, and GLM-5.3, a k-NN classifier matched 7 of 14 ox-alpha queries to GLM-5.3. Opus 5 came in second with 3 matches, not close enough to call it a coin flip.
A separate coding benchmark points the same direction. The stealth model scored 52 of 63 on the Rails benchmark, with 28.6% Rails API recall, numbers that track GLM-5.3's known range rather than a smaller model wearing its outputs.
The identity is one preview slot's story. Compression distance itself carries over to the next one. It turns "which model is actually behind this endpoint" into something testable from outside, with no cooperation from the host and no API access beyond the responses it already returns.
Each link below shares sources, entities, or timing with this story.
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
I've spent the last year assuming that if I wanted real agentic coding quality, I paid for a closed model. That assumption took a hit on June 1. MiniMax shipped M3 with a new sparse-attention architecture (they call it MSA) that handles up to 1M tokens at roughly 9x prefill an...
Zhipu's GLM-5.1 ranked third on Code Arena, jumping 90+ points over its predecessor GLM-5 and landing ahead of GPT-5.4 and Gemini 3.1 Pro. Two separate r/LocalLLaMA threads (491 upvotes and 233 upvotes) confirm this isn't just a benchmark curiosity. Practitioners are paying at...
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
This is the other half of the Fable 5 story, so read them together. While the best coding model in the world is uncallable, an open-weight one quietly posted frontier-adjacent numbers. Per Tom's Hardware, independent benchmarks for the MIT-licensed GLM-5.2 (744B params, 40B ac...
An open-weight model just beat every closed frontier model on the benchmark builders actually care about. Z.AI (formerly Zhipu AI) dropped GLM-5.1, a 754-billion parameter mixture-of-experts model with 40 billion active parameters. The SWE-Bench Pro score: 58.4%. That's above...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.