Fetching from the wire…
Research2026-08-13 · source-backed
arXiv 2608.11386 ran 11,700 repository-issue-fixing trajectories across six tool architectures with capabilities held roughly equal. Structured interfaces improved run-to-run consistency up to 4.7x, natural-language search tools raised relevant-file discovery over 11%, and Python CodeAct-style interfaces cut steps 41.6% and tokens 56.3%. Cognitive-scaffolding tools like reasoning logs showed almost no effect. The highest-leverage knob on your coding agent is the tool schema, not the system prompt.
Each link below shares sources, entities, or timing with this story.
arXiv 2607.08010 replaces the inference-time coding loop with a pipeline that collects execution traces, observes live backend schemas and values, synthesizes candidate tools, and repairs them against labeled cases. Runtime calls the compiled tool instead of re-deriving it. Th...
OpenAI Devs announced on August 26 that WebMCP works in the ChatGPT desktop app's built-in browser and in ChatGPT Sites, so ChatGPT and Codex can call a site's declared tools directly. WebMCP is an experimental web standard adding navigator.modelContext to the browser, letting...
Nearly everyone wraps their agent instructions in XML tags. The vendor docs implied it helped, so it propagated, and now <instructions> and <rules> blocks are the house style of the entire industry. A deployed tender-response system measured it and found the formatting rule is...
1. Harden CI/CD Pipelines Against PromptPwnd AI Injection | Intermediate Aikido Security disclosed "PromptPwnd" — five Fortune 500 companies confirmed affected by AI agent injection in GitHub Actions. 1. Audit all .github/workflows/ for user-controlled input (github.event.issu...
A paper from Stanovsky's group (arXiv:2606.16576) tests whether LLM agents can uncover a hidden deterministic finite automaton through membership and equivalence queries. Performance drops sharply as the automaton grows, and trajectory analysis exposes recurring failures in qu...
ZeroDayBench (2603.02297) — GPT-5.2, Claude Sonnet 4.5, and Grok 4.1 all fail at autonomous zero-day vulnerability discovery. Reality check: the CyberStrikeAI threat is automation of *known* exploits, not novel vulnerability discovery. ICLR 2026 Workshop. tau-Knowledge (2603.0...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.