Fetching from the wire…
Research2026-08-20 · source-backed
Using CoderForge-Preview, described as the largest open dataset of coding agent trajectories, ensemble methods with SHAP attribution predict agent success before any run. Dominant drivers are patch fragmentation (how many places a fix has to touch) and repository scale. Prompt linguistic features only matter in the mid-difficulty band. (arXiv 2608.18280) Fragmentation over scale is the useful bit: a 40-line change across eight files is harder for an agent than a 400-line change in one.
Each link below shares sources, entities, or timing with this story.
Guardrails evaluate one session at a time; real adversaries spread attacks across independent agents and runtimes so each local defense sees a sparse fragment (arXiv 2607.18826). Asynchronous Attribution Fingerprint Vectors score campaign similarity from tool-use patterns, tim...
Logistic-regression probes on a coding agent's hidden states can decode whether code will parse and pass tests at AUC up to 0.83, and those internal representations run ahead of the agent's own edits, predicting outcomes as much as 25 steps in advance (arXiv). The authors call...
On a verifiable protein-function characterization task routed across tools, model choice swamped federation topology, RL-versus-LLM harness, and prompt expertise: Opus at roughly 92 to 94%, o4-mini at 40 to 50%. Federation across institutional boundaries cost almost nothing (a...
Agents share transport and can call each other's tools but have no way to reconcile a fact phrased two ways (arXiv 2608.16357). Every incoming claim passes a five-outcome procedure (insert, merge, relate, conflict, reject) decided from scoped claim-key identity, embedding simi...
arXiv 2608.10314 had two LLM snapshots translate five theoretical accounts into code under structured-contract versus prose formats, producing 320 programs. Both primary hypotheses returned NOT_SUPPORTED, only 19 of 108 criterion evaluations passed, and cross-model format iden...
Scoring each retrieved chunk and dropping failures assumes one chunk is a sufficient premise; multi-hop questions are built so none is. Entailment scoring reaches 0.643/0.523/0.560 AUC on HotpotQA, 2Wiki, and MuSiQue against 0.951 on single-hop SQuAD, and per-chunk gating was...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.