Fetching from the wire…
Research2026-06-29 · source-backed
The June 29 Download flags the structural gap between what benchmarks measure and what actually matters. Read it right after the GLM-5.2 numbers above. A 39% F1 on IDOR detection is a real signal, but a benchmark score is not the same as the model being good at your task. Ground your evals in the work, not the leaderboard.
Each link below shares sources, entities, or timing with this story.
MIT Technology Review's Download covers a transmission line fight that could reshape New York's grid alongside escalating US threats against Chinese AI (MIT Tech Review). Transmission, not generation and not silicon, is increasingly what determines data center siting. AI expan...
codenotch is a Swift menu-bar app at 1,011 stars, created September 5, that reads Claude Code's /usage output (falling back to the OAuth token in the login keychain), Cursor's local SQLite session state, Codex's ChatGPT usage endpoint for 5-hour and weekly limits, plus Antigra...
A September 1 analysis rebuilds Artificial Analysis's chart on a linear rather than logarithmic cost axis and prices models at what third-party providers actually charge. The spread is roughly 250x top to bottom: Fable 5.1 at $3.69 per task for intelligence 66, GLM-5.3-Flash a...
Unisound's U2 is a 266B-total / 10B-active MoE tuned for agents, citing 72.2% SWE-bench Verified at $0.15/$0.30 per 1M tokens, while GLM-5.2 is getting named the strongest open-weight coding model across July roundups. Both are roundup-sourced, so verify the benchmarks against...
The project (pure C, Apache-2.0, 71 stars) lazy-fetches only the bytes an inference touches and caches them locally, sending 4 KB activations to peers holding the relevant experts rather than transferring expert weights. Local and remote paths share identical code to guarantee...
Snowflake's HybridDeepResearch supplies 380 tool-dependent tasks grounded in LiveSQLBench-Base-Lite databases plus public web corpora, testing whether an agent preserves constraints while moving evidence between systems. GLM-5.2, Claude Sonnet 4.6 and GPT-5 all cluster around...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.