Fetching from the wire…
Agents2026-07-10 · source-backed
arXiv 2607.08010 replaces the inference-time coding loop with a pipeline that collects execution traces, observes live backend schemas and values, synthesizes candidate tools, and repairs them against labeled cases. Runtime calls the compiled tool instead of re-deriving it. This is the practical middle ground between hardcoded pipelines and fully dynamic code generation, and it pairs naturally with programmatic tool calling: compile the tool once, let the model orchestrate it in JS.
Each link below shares sources, entities, or timing with this story.
Introduces temporal causal diagnostics to distinguish legitimate task execution from injected manipulation in multi-turn agent interactions, plus context purification to neutralize poisoned content. Directly applicable to anyone building agents that call external tools. arXiv...
OpenAI Devs announced on August 26 that WebMCP works in the ChatGPT desktop app's built-in browser and in ChatGPT Sites, so ChatGPT and Codex can call a site's declared tools directly. WebMCP is an experimental web standard adding navigator.modelContext to the browser, letting...
arXiv 2608.11386 ran 11,700 repository-issue-fixing trajectories across six tool architectures with capabilities held roughly equal. Structured interfaces improved run-to-run consistency up to 4.7x, natural-language search tools raised relevant-file discovery over 11%, and Pyt...
1. Harden CI/CD Pipelines Against PromptPwnd AI Injection | Intermediate Aikido Security disclosed "PromptPwnd" — five Fortune 500 companies confirmed affected by AI agent injection in GitHub Actions. 1. Audit all .github/workflows/ for user-controlled input (github.event.issu...
arXiv 2608.06370 evaluated models emitting code that calls tools against JSON-schema tool calling on BFCL v4. PTC matched or exceeded the baseline in 11 of 14 models, with the GPT-5.6 family up 10.6%, and held stable under parallel execution in 13 of 14. Under context degradat...
arXiv:2606.07889 names a failure mode where a coding agent holds information that should change its behavior, states that information out loud, and then acts against it anyway. The authors propose detecting this in execution trajectories as a pre-failure signal. For anyone run...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.