Fetching from the wire…
Public story · 2026-08-31 · high
The preview weighs 1.56TB and shows its reasoning live, weighing a helmet and sunglasses before cutting both from an SVG.
Why now: Simon Willison tested the Hy4 preview and posted his notes on August 29, 2026.
Tencent's Hy4 preview packs 770 billion total parameters, 49 billion active, and a 1 million token context window, per Simon Willison's testing. The full weights come to 1.56TB. That size puts self-hosting out of reach for most local setups.
Hy3, the prior version, ran 295 billion total parameters with 21 billion active and a 256K context window. Hy4 also ships two reasoning effort levels, with "high" as the default and a "no_think" mode for skipping the extra deliberation.
Willison's test was an SVG generation task. The reasoning trace showed the model weighing specific details, considering a helmet, considering sunglasses, then rejecting both. It read in slightly truncated English, more like notes than full sentences.
Willison's writeup doesn't include a benchmark comparison against Hy3 or against other large models at similar active-parameter counts. He reported what he saw running it himself. He didn't say whether the extra size or the larger context window improves output quality beyond that one SVG test.
Each link below shares sources, entities, or timing with this story.
78 layers where layer one is dense FFN and the other 77 are MoE, each with 256 routed experts and 1 shared expert, top-8 routing per token, plus a native 10B MTP layer (0.7B activated) built in for speculative decoding. FP8 and base variants released together on August 28; the...
Satya Nadella said companies routing everything through a single proprietary lab may not survive. His argument: you hand that lab your most sensitive business context, and the lab can turn it against you as a competitor. His prescription is an orchestration layer — keep the ha...
Hy4 preview, released August 28: 770B total parameters, 49B active, over 1M token context, open-sourced and simultaneously on Tencent Cloud TokenHub and OpenRouter at $0.834 per million input tokens, $2.501 per million output, $0.042 per million cached. In Tencent's own evalua...
Released August 28 with 78 layers, 77 of them MoE with 256 routed plus one shared expert and top-8 routing, plus a native 10B MTP layer for speculative decoding (GitHub). The attention stack uses Gated DeepSeek Sparse Attention with IndexCache for cross-layer sparse index reus...
The August 16 upgrade to his markdown renderer detects whether an SVG contains SMIL or CSS animation, guesses the loop duration, renders the frames, then loads ffmpeg.wasm to compile them into a downloadable MP4 entirely client-side (simonwillison.net). No server, no upload. T...
Simon Willison doesn't hand out superlatives. So when he writes that Z.ai's GLM-5.2 is "probably the most powerful text-only open weights LLM," that's worth stopping for. His June 17 evaluation walks through a 753B-parameter Mixture-of-Experts model with 40B active params, a 1...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.