Fetching from the wire…
Models2026-08-28 · source-backed
One file handles tokenizer, transformer, KV cache, sampling and CPU kernels with no external library doing the interesting parts, producing about a 5.0 GB model file (GitHub). Weights are int8 with FP16 scales, linear-layer inputs dynamically quantized to int8, other activations float32, validated against the Hugging Face reference running Google's unquantized QAT checkpoint in BF16. The author confirms the demo is real time, CPU only. Read it if you want to know what a competent inference runtime is actually doing.
Each link below shares sources, entities, or timing with this story.
Danish Foundation Models trained it from scratch on 161 datasets. Across 20 benchmarks spanning English, math and code, and Danish, it beats the original HRM-Text 1B, sets a new Danish state of the art, and competes with Qwen 3.5 4B and Gemma 4 E2B. Weights are on Hugging Face...
Google's HF org lists diffusiongemma-26B-A4B-it (~4B active), an image-text-to-text Gemma member that's diffusion-style rather than purely autoregressive (Hugging Face). No detailed announcement yet, which is why I'm flagging it low. But a diffusion approach inside the Gemma o...
1. Use claude agents --json to build session dashboards. Claude Code v2.1.145 outputs all live agent sessions as structured JSON with status, model, elapsed time, and parent relationships. Pipe it into a tmux status bar widget or session picker script for switching between bac...
A spec is a press release until someone who didn't write it implements it. GitHub made Agent Plugins 1.0 generally available on August 12 across VS Code, Copilot CLI, the Copilot SDK, and the Copilot app on all plans. The spec, published August 6, was co-authored by AWS, Anysp...
google-research/timesfm gained 326 stars, second on the all-language board. google/timesfm-3.0-pytorch was created August 24 and shows 257 likes against 0 downloads (Hugging Face). Nine days of likes with no downloads means people are bookmarking, not running. Useful calibrati...
modelprint runs 9 infrastructure probes against any OpenAI-compatible endpoint from a static page with no server, keys never leaving the tab. Its day-one run against 12 candidates scored stealth/ox-alpha at 6 of 9 probes and 4 of 4 normalized tokenizer counts matching z-ai/glm...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.