Fetching from the wire…
Research2026-08-27 · source-backed
The authors extract a steering direction from the model's existing tool-use preference signal and apply it at inference, producing monotonic control over how often the agent reaches for a tool while keeping invocations valid (arXiv 2608.25198). Open-domain QA accuracy with live tool execution nearly doubled, from 0.29 to 0.56, by tuning the rate rather than the prompt. It generalizes to unseen tools and works across dense, MoE and multimodal architectures. That makes tool-call frequency a deployment knob you set per environment instead of a paragraph you keep rewriting in the system prompt.
Each link below shares sources, entities, or timing with this story.
NVIDIA released Nemotron-Cascade 2 — a 30B MoE model activating only 3B parameters per token, trained with Cascade RL and multi-domain on-policy distillation. Claims best-in-class reasoning among open models at its efficiency tier with strong agentic task performance. The Casc...
A 15-author Huawei team argues kernel-generation benchmarks are almost entirely CUDA and Triton, leaving less-documented hardware with no shared yardstick (arXiv 2607.20518). CANN Bench covers 53 operators and 1,060 test cases in four difficulty tiers, from elementwise primiti...
A systematic negative rounding error in non-uniform FP4 (E2M1) formats compounds across layers, degrading training (arXiv:2606.20381). The proposed UFP4 recipe trains on uniform E1M2/INT4 grids with a Random Hadamard Transform on all three GEMMs, and shows lower loss degradati...
Load Hijack modifies nothing but router weights in a checkpoint. When a private trigger appears, token-to-expert assignment concentrates on experts co-located on a single GPU, making it a straggler while peers idle (arXiv 2608.10614). Across three MoE families and four corpora...
An r/LocalLLaMA post reports a working deployment on 80x RTX 5090 connected over 25 gigabit Ethernet rather than NVLink or InfiniBand — roughly 2.5TB aggregate VRAM against a ~594GB MXFP4 weight file, surplus absorbed by activation and KV overhead. Expert-parallel MoE over com...
A llama.cpp fork by fewtarius aimed at AMD APUs, iGPUs and handhelds adds a persistent SSD-backed KV cache with hot/warm/cold tiering. On an Ayaneo Flip KB running Qwen3.6-35B over a 15,700-token prompt, cold TTFT was 143.1 seconds versus 0.99 seconds warm, a 144.5x speedup. A...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.