Fetching from the wire…
Public story · 2026-08-31 · high
The paper says matching those gains with linear attention would require training from scratch or heavy post-training.
Why now: The comparison arrives as linear-attention retrofits keep getting pitched as a memory-saving swap for long-context models.
Sliding window attention with sinks beats post-trained linear attention on long-context retrieval. It scores 2 to 10 times higher on the Needle-in-a-Haystack and BABILong benchmarks, per a new comparison of attention mechanisms that tested both approaches across multiple LLMs and downstream tasks.
That gap matters for anyone weighing how to cut inference memory costs on long-context models. Linear attention has become a popular retrofit for that exact problem: take an existing model, post-train it into a linear-attention variant, and save memory without retraining from scratch.
The comparison ran the baseline retrofit papers tend to skip. It tested SWA with sinks directly against post-trained linear models on the same tasks, not just the two retrieval benchmarks but the broader set of downstream tasks too. SWA matched or beat linear attention across the board.
The paper's recommendation is blunt. Switch to SWA for the memory savings. Matching its retrieval scores with a linear model would take training from scratch or heavy post-training, the paper says, not a light retrofit.
It's a negative result for a popular idea. Papers proposing linear-attention retrofits don't tend to run this comparison, which is exactly why it's useful now.
Each link below shares sources, entities, or timing with this story.
Responsible Statecraft reported Aug 17 that the "Hanover Institute for Public Policy" is a front created by Piro, Inc. under a $900,000 contract from the Israeli Government Advertising Agency, subcontracted through Havas Media. Piro's own site markets the practice as "AI Story...
A GitHub Issue. No code, no credentials, no access. Just a paragraph of English that tells an AI agent to copy your private repo into a public comment. That's GitLost, and it works whether the agent runs on Copilot, Claude, Gemini, or Codex. (Noma Security) Noma Security discl...
LLMs are brittle to renamed nodes and reworded formulations in graph reasoning, and the standard fix is throwing a multi-agent system at the parsing failures. GRAIN is a single RL-trained agent modeling reasoning as semantic parsing plus tool execution, rewarded by a Structure...
ES achieves broader reasoning coverage, with verifier-projected Jensen-Shannon diversity across the ES population theoretically tied to higher Pass@K, and empirically improves Pass@1 while reaching higher Pass@K where GRPO collapses entropy. The proposed sequential GRPO-then-E...
The serving system targets the practical inference-cost wall you hit when agents produce very long outputs and dense attention becomes the bottleneck (arXiv). As agent runs get longer, this is the kind of infra that decides whether serving them at scale is affordable. Directly...
A prespecified randomized audit ran seven models over 3,024 choice sets, three personas, nine paraphrases and nine arms for 40,068 scored responses (arXiv 2608.14399). Reputation dominates, with a 3.9 to 4.7 rating raising choice probability 31.4 points. But demographic parity...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.