Fetching from the wire…
Research2026-08-12 · source-backed
The argument is that token uncertainty shows up not only in output-distribution breadth but in whether a confident prediction is fragile under perturbation of its attention pathways (arXiv 2608.11138). It's training-free: mask attention heads, measure BALD mutual information among resulting subnetworks through a semantic-agreement kernel. Single-response variant ties or beats Semantic Entropy on 10 of 12 grounded benchmark-backbone settings. The honest limit the authors state plainly: on parametric QA all variants revert to or below the zero-cost MSP baseline. This helps for retrieval-grounded answers, not closed-book ones.
Each link below shares sources, entities, or timing with this story.
Learned KV eviction has a soft-to-hard mismatch: training uses differentiable gates that attenuate contributions, but inference only saves memory when entries are physically removed (arXiv 2608.23296). A controlled 2x2x2 over attention type, learned gating and positional encod...
Diverse Hypothesis Deliberation caches five independently generated messages per problem, then hides and reveals each to the same downstream integrator to measure marginal contribution (arXiv 2608.14375). Across five math and science benchmarks and two model families, wrong-bu...
arXiv 2607.27951 separates the capability a model releases from the evidence it has about downstream use, and shows that when that evidence is copyable (a request, a persona, an interaction history an attacker can imitate) there's an exact worst-case floor on attacker assistan...
Sergey Rodionov's paper tests four Codex-based agent variants to isolate what actually drives performance. Verification (simplification plus exact observation reproduction) ranked highest in every setting, but at substantially higher cost. The textual baseline beat the executa...
Attnlocate (arXiv 2608.24022) aggregates attention across heads and layers into a token-level feature space, then runs a 1-D U-Net with an anchor-free detection head to find the traces behavior-guiding instructions leave behind, adjudicating the tool call based on the authorit...
The study extracted 130 clean atomic state transitions from 707 real issues in SWE-bench Lite and Verified. Plain RAG scored 0.57-0.59 answer accuracy; an LLM reranker didn't help and added latency, about 18 seconds against 2.1. A (subject, relation, object) supersession memor...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.