Fetching from the wire…
Agents2026-08-30 · source-backed
LLMs are brittle to renamed nodes and reworded formulations in graph reasoning, and the standard fix is throwing a multi-agent system at the parsing failures. GRAIN is a single RL-trained agent modeling reasoning as semantic parsing plus tool execution, rewarded by a Structure Invariance Reward that validates extracted intermediate graphs against ground-truth topology. It halves the out-of-distribution gap of SFT models from 15.77% to 7.80%. (arXiv 2608.27142) Second result this issue where a trained single agent beat an orchestrated multi-agent baseline.
Each link below shares sources, entities, or timing with this story.
This comparison ran the baseline the retrofitted-linear-attention literature skipped. Across multiple LLMs and downstream tasks SWA with sinks matches or beats post-trained linear attention, and on Needle-in-a-Haystack and BABILong it scores 2 to 10 times higher. The recommend...
LangChoiceBench covers 28 projects across seven software areas where Python is a poor default, run against 25 LLMs. Python stays heavily over-selected, recommendation-implementation consistency is low, and smaller open-weight models show stronger bias. Analysis of 9,826 reason...
arXiv 2607.23982 adapts Holmström's team moral-hazard model into a game where an agent can keep an immediate local reward or pay a query cost to surface a hidden safety fact that mainly helps another agent's downstream decision. Base behavior splits into two failure modes: pre...
A new paper on "Contagion Networks" (arXiv:2606.20493) shows that when LLMs serve as evaluators inside multi-agent systems, their systematic biases propagate through the network rather than staying local. A single biased judge can contaminate downstream agent decisions. If you...
Studdiford and Lupyan tested human participants and 25 LLMs on common-sense causal reasoning and found shared, predictable error patterns triggered by irrelevant prompt details (arXiv). They localized the attention heads driving it. The uncomfortable implication: the gap betwe...
SynChain uses persistence-aware directed SFT to make a computer-use agent produce artifacts that pass vetting while hiding malicious influence in structural redundancies, surviving internal state updates and reactivating in a later workflow with no new external input. Tested a...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.