Fetching from the wire…
Research2026-08-06 · source-backed
arXiv 2608.04893 tests the "exchanged latent thoughts" claim by replacing the relayed cache with deranged, zeroed and moment-matched random counterparts. The claim holds only when the receiver genuinely needs the sender's private information (100% vs 23-25%, replicated across three families and five checkpoints). Otherwise a pre-registered five-seed protocol establishes equivalence within 2.8 points under Holm-corrected TOST. In one cell, zeroing the relay costs 14.7 points while a mismatched cache costs 0.4. Among released systems: LatentMAS's relay hits ceiling, KVComm's layer subset is partial, C2C's projector shows no detected example-specific transfer.
Each link below shares sources, entities, or timing with this story.
Mehan and Saluja audited 200 open-source Python microservice projects. Explicit retry logic is detected in 11.5%, though their own false-negative audit puts true prevalence near 41%. Among detected projects 60.9% have at least one configuration with no backoff, and exactly one...
arXiv 2607.25886 isolates data-centric research capability by fixing the entire post-training stack so only the agent's data strategy varies. Four frontier agents across six benchmarks. Among searches that continued past the best observed score, 78.26% ended on a lower-scoring...
arXiv 2608.09624 separates harmful intent (a prompt property) from jailbreak success (an outcome from a specific model, decoder, and judge). On Llama, wrapping a prompt raises harmful generation from 0.05 to 0.27 while harmful-intent AUROC falls from 0.936 to 0.803. Attacks ge...
arXiv 2607.12227 (Wang et al., incl. Hajishirzi, Tsvetkov, Dasigi) finds two methodological holes in the self-improving-agent literature: methods are never compared against simpler baselines at matched compute budgets, and final performance gets reported on the same public ben...
Stanford's Denisov-Blanch group built a maturity model for AI adoption scored entirely from artifacts already in version control, applied it to 441 repositories, and found something I've been assuming without evidence. RAMP is a four-level model derived only from committed AI...
Every retrieval pipeline I've built follows the same instinct: rank, threshold, pass only the top hits. Noise is bad. Precision is good. A controlled study says that instinct costs you accuracy (arXiv 2608.17188). 2,420 trials, 11 model configurations, 661 anonymized workplace...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.