Fetching from the wire…
Security2026-08-20 · source-backed
Task-Conditioned Least-Privilege Learning trained Qwen3.5-4B on 1,500 terminal and MCP tasks to choose an authority level that fits the task, audited by deterministic verifiers across six risk dimensions. Over 2,896 episodes on 500 held-out tasks, excess-authority errors fell from 4.56% to 0.79%. (arXiv 2608.18351) The authors are refreshingly clear that learned restraint sits on top of permission gates and sandboxes, it does not replace them. Which is the right framing: a model that usually asks for less is not a security boundary, it's a way to reduce how often your actual boundary has to say no.
Each link below shares sources, entities, or timing with this story.
Thibault Sottiaux at OpenAI published an investigation into "a handful of reports where GPT-5.6 unexpectedly deleted files," finding it happens most commonly when full access mode is enabled in Codex. Simon Willison relayed it. A frontier lab publishing a first-party post-mort...
This one annoyed me, because I've been running the losing pattern. SWE-QA (arXiv 2608.01507) compares the sub-agent grep pattern that Claude Code, Codex and Antigravity all ship by default against a pre-built semantic index over the same repository. Semantic search answered 65...
OpenAI Devs announced on August 26 that WebMCP works in the ChatGPT desktop app's built-in browser and in ChatGPT Sites, so ChatGPT and Codex can call a site's declared tools directly. WebMCP is an experimental web standard adding navigator.modelContext to the browser, letting...
xAI launched Grok Build on May 14. With that, every major AI lab now ships a coding agent that lives in your terminal. The competition isn't "can we build one" anymore. That question is settled. The lineup: Anthropic has Claude Code. OpenAI has Codex CLI. Google has Gemini CLI...
| # | Skill | Domain | Difficulty | |---|-------|--------|------------| | 1 | Claude Code /simplify + /batch — three-agent parallel review + codebase migrations | vibe-coding | intermediate | | 2 | Pipelock agent firewall — 9-layer DLP + MCP scanning inline proxy | agent-secur...
Nemotron-Terminal (arXiv:2602.21193, 66 HF upvotes) — First systematic study of data engineering for terminal/CLI agents. Terminal-Task-Gen pipeline with Dockerized environment interaction. Qwen3-initialized 8B model goes from 2.5% to 13.0% on Terminal-Bench 2.0. All checkpoin...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.