Fetching from the wire…
Research2026-08-15 · source-backed
The standard protocol ablates a latent and measures effect at the token where it fires hardest, but that token is chosen by the dictionary under evaluation. Two dictionaries get compared at different places. Training six autoencoders from one initialization showed 7.6% and 11.9% of apparent inter-dictionary variance collapses to near zero once position is held fixed. Across a sixteenfold range of corpus sizes, dictionaries agreed less about where to measure. arXiv 2608.13337 The problem grows with scale, and the paper audits five published papers against a corrected protocol.
Each link below shares sources, entities, or timing with this story.
The diagnosis in this paper is better than the fix, and the fix is very good. Recurrent memory agents fail at long context, but not for the reason most people assume. The bottleneck isn't capture. It's retention. Retention falls below 30% at 896K tokens because every consolida...
Prior work found individual parameters whose removal collapses LLM performance by orders of magnitude. arXiv 2607.08733 shows the effect isn't universal across models, then tests the obvious corollary that Super Weight-aware training should work. It doesn't. Training 100 to 8,...
arXiv 2608.12253 shows the standard practice of training a policy against a single LLM simulating the user fails because the simulator is itself mode-collapsed, so the policy learns to exploit its dominant mode. Verbalized Sampling recovers up to 9% held-out success; Populatio...
RETRACE has a verifier infer what problem the patch appears to solve using only the patch and trajectory, then compares that inference against the real issue. Training-free, lifted Pass@1 by 7.0% and 3.6% on mini-SWE-agent over SWE-bench Verified. The information-hiding trick...
arXiv 2608.13030 points out existing agent protocols specify message exchange but not how an agent proves identity, authorization, advertised capabilities, or accountability after delegation. It adds Persistent Identity, Discovery, Trust Negotiation and Accountability layers v...
SMITH (arXiv 2608.24571) points at a real gap in existing tool-creation systems: they prompt a frozen LLM at inference time, so the model writing a schema gets no signal about whether it can invoke that schema. SMITH alternates build rollouts and use rollouts inside one RL pol...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.