Fetching from the wire…
Public story · 2026-03-19 · source-backed
ArXiv 2603.17942 demonstrates that standard next-token models exhibit latent multi-token prediction capabilities extractable via lightweight embedding-space probes, with no additional training required. Inference speedups match explicitly MTP-trained models. Existing deployed models can be accelerated with MTP probes as a zero-cost optimization — challenging the assumption that dedicated training objectives are required. arXiv
Each link below shares sources, entities, or timing with this story.
Existing neuron-level defenses stay always-on and perturb every benign request (arXiv 2608.14392). Tripwire identifies safety-specific neurons through per-neuron hypothesis tests under false-discovery-rate control plus a utility-specificity filter, then clamps them to harmful-...
Google released open-source Multi-Token Prediction (MTP) drafters for the Gemma 4 model family. The concept: pair a heavy target model (Gemma 4 31B) with a lightweight drafter that predicts several future tokens in parallel. The target model verifies the predictions in a singl...
arXiv 2609.20734 shows a pretrained model's decoding states already carry information predictive of whether a global read will help, before the read happens. Training only a small recall head to invoke global attention selectively, with pretrained weights untouched and the ful...
The target is attacks where a harmful objective is split across individually plausible requests and tool calls, so the harm is only visible in the accumulated trajectory. Existing defenses either pay for auxiliary online reasoning or judge actions after generation, which ties...
SMITH (arXiv 2608.24571) points at a real gap in existing tool-creation systems: they prompt a frozen LLM at inference time, so the model writing a schema gets no signal about whether it can invoke that schema. SMITH alternates build rollouts and use rollouts inside one RL pol...
Prior work found individual parameters whose removal collapses LLM performance by orders of magnitude. arXiv 2607.08733 shows the effect isn't universal across models, then tests the obvious corollary that Super Weight-aware training should work. It doesn't. Training 100 to 8,...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.