Fetching from the wire…
Public story · 2026-08-17 · high
The method beat Gemini-2.5-Pro on ProofBench and repeated across Qwen3, GLM-4.7 and Gemma-4, per the arXiv paper.
Why now: The technique is new as of August 17, 2026, per the arXiv listing.
SimpleOPD distills language models across tokenizer boundaries by working in shared text space, lifting Intern-S2-Preview 21.2 points on ProofBench, per a new arXiv paper.
Teacher and student models rarely share a tokenizer. That gap has forced teams to pick a compatible but weaker teacher instead of whichever model scores highest.
The method fixes this by distilling in shared text space instead of matching token IDs. It adds a student reference KL loss and masks the advantages of special termination tokens, which stops the student's generations from running long.
Intern-S2-Preview reached 55.2 on ProofBench with the technique, past Gemini-2.5-Pro's score. The paper reports the same gains across Qwen3, Qwen3.5, Intern-S2, GLM-4.7 and Gemma-4, and the improvement generalizes to two other benchmarks, HLE and HiPhO.
The paper doesn't report what the extra KL loss costs in training time or compute. For a technique meant to widen teacher choice, that's the number builders adopting it will want next.
The tokenizer match that used to gate which teacher a team picked no longer applies. Benchmark score becomes the only reason left to pick one model over another for distillation.
Each link below shares sources, entities, or timing with this story.
The models are good. The license is the real story. Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. Th...
Zhipu's GLM-5.1 ranked third on Code Arena, jumping 90+ points over its predecessor GLM-5 and landing ahead of GPT-5.4 and Gemini 3.1 Pro. Two separate r/LocalLLaMA threads (491 upvotes and 233 upvotes) confirm this isn't just a benchmark curiosity. Practitioners are paying at...
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
Zhong, Raghunathan, Laidlaw and Steinhardt fed 280 identities through Claude Code across four tasks. Against recognized safety researchers versus general users, Claude dropped behavioral confidence 1.4pp, increased reasoning usage 4.0pp and graded 0.11 points harder. Being tol...
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
40. HBR — AI intensifies work 41. Import AI #441 42. Simon Willison — AI hit piece 43. Cameron Wolfe — GRPO++ 44. CISPA — Moltbook study 45. arXiv — SkillRL 46. Google DeepMind — Gemini 3 Deep Think 47. OpenAI — Codex-Spark 48. Zhipu — GLM-5 ---
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.