Fetching from the wire…
Agents2026-08-28 · source-backed
Built from 542 quality-controlled factual questions into 6,504 episodes with tool returns of known correctness, across five open-weight 7-9B models (arXiv 2608.26295). Models follow a correct tool 86.0-93.1% of the time and repeat the tool return in 78.4-86.0% of cases where both the tool and the model are wrong. No cross-model ordering survives three instruction-wording variants with content held fixed, so tool-trust behavior is unstable to prompt phrasing. Practical read: a poisoned tool output beats the model's own correct knowledge roughly nine times in ten.
Each link below shares sources, entities, or timing with this story.
This one annoyed me, because I've been running the losing pattern. SWE-QA (arXiv 2608.01507) compares the sub-agent grep pattern that Claude Code, Codex and Antigravity all ship by default against a pre-built semantic index over the same repository. Semantic search answered 65...
Models navigate to the correct file for 92%+ of required deletions but cut the exact target line only 52% of the time, and 29% of passing patches wrap dead code in a conditional instead of removing it. Grep the diff for newly added if guards around code the task said to delete...
HANDBOOK.md is a benchmark for whether standing instructions actually constrain an agent across extended tool-use runs. Not whether the model reads your policy file. Whether it still obeys it forty tool calls deep. 65 tasks pairing expert-written SOPs of 20 to 124 pages across...
arXiv 2608.06370 evaluated models emitting code that calls tools against JSON-schema tool calling on BFCL v4. PTC matched or exceeded the baseline in 11 of 14 models, with the GPT-5.6 family up 10.6%, and held stable under parallel execution in 13 of 14. Under context degradat...
Google research showing CoT reasoning substantially expands LLM parametric knowledge recall — unlocking correct answers unreachable via direct prompting, even for simple factual questions. Practical implication: always-on reasoning may be worth the cost for knowledge-intensive...
TRAPSBench is a procedurally generated video benchmark of 1,404 matched physics pairs where one targeted change makes the outcome undeterminable from the visuals, scored by a Penalized Epistemic Calibration Score demanding correct answers when knowable and abstention when not....
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.