Fetching from the wire…
Research2026-08-21 · source-backed
A synthetic benchmark constructs conflicts where exactly one evidence source matches ground truth, independently varying modality, recency, stated reliability, and provenance. Across open-weight instruction-tuned models the arbitration is systematic: distinct text-versus-number preferences, temporal recency weighted more consistently than explicit reliability cues, and over-trust of external forecasts contradicting direct in-context evidence. arXiv An out-of-date but recent-looking tool response quietly outranks a correct fact you put in the prompt. Labeling your sources "authoritative" does less than you think.
Each link below shares sources, entities, or timing with this story.
Here's a finding that goes against the thing everyone assumes. We tell ourselves that as base models get more capable, agents built on them will get more discerning about their tools, second-guessing bad outputs, catching errors, adding reasoning on top. A new study says the o...
Two thirds. Not two thirds of a contrived jailbreak set. Two thirds of realistic malicious issue requests, against the exact three tools most of the people reading this run daily. Ankur Singh, Jinqiu Yang, and Tse-Hsun Chen built IssueTrojanBench across four attack categories...
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
This paper separates interpreting a source from aggregating interpretations, proposing a four-field evidence tuple as the interface, and names a failure mode hitting any additive scoring system (arXiv 2608.14509). Thresholding a sum of unnormalized weights is posterior thresho...
A symmetry-based framework built 550 conflict instances across 11 function families and four representation types with explicit built-in contradictions. Resolution is systematic, not random, with a consistent ordering: Formal ≈ Naturalized Formal > Pure Natural Language > Inpu...
Every coding harness I've built, including the one that produces this newsletter, has some version of "if it fails, try again." That instinct is wrong, and there's now a study with the seed count to prove it. "Looping Is Not Reliability" (arXiv 2607.24604, July 27) runs a seal...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.