Fetching from the wire…
Agents2026-08-18 · source-backed
Agents share transport and can call each other's tools but have no way to reconcile a fact phrased two ways (arXiv 2608.16357). Every incoming claim passes a five-outcome procedure (insert, merge, relate, conflict, reject) decided from scoped claim-key identity, embedding similarity, and an NLI verdict, mutating state only through one authenticated Patch. 0.968 AUC, 0.013 false-merge rate. The status CRDT reconverges in 30/30 partition-heal trials where last-writer-wins manages 11/30, and semantic routing delivers ~3x fewer messages at matched recall. Contradictions get preserved rather than silently overwritten, which is the part most memory systems get wrong.
Each link below shares sources, entities, or timing with this story.
arXiv 2608.23740 exposes file-level claim, status and broadcast as MCP tools over a shared filesystem, motivated by the observation that a single agent abandons up to half of hard tasks with a one-file stub-and-exit. Across five frontier coding CLIs on four backend tasks, two-...
Guardrails evaluate one session at a time; real adversaries spread attacks across independent agents and runtimes so each local defense sees a sparse fragment (arXiv 2607.18826). Asynchronous Attribution Fingerprint Vectors score campaign similarity from tool-use patterns, tim...
Logistic-regression probes on a coding agent's hidden states can decode whether code will parse and pass tests at AUC up to 0.83, and those internal representations run ahead of the agent's own edits, predicting outcomes as much as 25 steps in advance (arXiv). The authors call...
"Memory in the Loop" (arXiv:2607.05690) moves memory read/write inside the agent's per-step loop, viable only with an in-process store answering in ~100µs. The behavioral number is the story: redundant actions were 0.0 of 12 at in-process speed but 7.2 of 12 at a 110ms cloud r...
arXiv 2608.10314 had two LLM snapshots translate five theoretical accounts into code under structured-contract versus prose formats, producing 320 programs. Both primary hypotheses returned NOT_SUPPORTED, only 19 of 108 criterion evaluations passed, and cross-model format iden...
Scoring each retrieved chunk and dropping failures assumes one chunk is a sufficient premise; multi-hop questions are built so none is. Entailment scoring reaches 0.643/0.523/0.560 AUC on HotpotQA, 2Wiki, and MuSiQue against 0.951 on single-hop SQuAD, and per-chunk gating was...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.