Fetching from the wire…
Skills2026-08-09 · source-backed
Clone Boundary-Bench, point it at your actual nftables/Landlock/setpriv configuration, and run the 89 Terminal-Bench 2.1 tasks. Vendors' leaderboard numbers come from unrestricted environments; the leader flips under hardening and costs inflate up to 167%. A day of work gets you a number that survives contact with your EDR.
Each link below shares sources, entities, or timing with this story.
Every coding agent leaderboard number you've seen was produced in conditions your security team would reject on sight. Researchers from Accomplish AI and NYU open-sourced Boundary-Bench on August 5. The setup: 12 frontier agent harnesses, roughly 10,000 runs, 89 Terminal-Bench...
Terminal-Bench 2.1 results (entries dated June 17) put Codex CLI on GPT-5.5 first at 83.4%, Claude Code on Fable 5 second at 83.1%, and Claude Code on Opus 4.8 at 78.9%. The asterisk matters more than the ranking: Fable 5 and Mythos 5 have been export-suspended since June 12,...
Cline released @cline/sdk on May 13, an open-source TypeScript agent runtime that powers their CLI, VS Code, and JetBrains extensions. Running claude-opus-4.7, Cline CLI scores 74.2% on Terminal-Bench 2.0. Claude Code on the same model: 69.4%. Same model. Different harness. Al...
Claude Code 2.1.229 shipped a config flag most people will scroll past: CLAUDE_CODE_WORKFLOW_PREFIX_STAGGER_MS. What it does is delay the launch of sibling agents that share a prompt prefix, so the second through Nth agents read the warm cache instead of each writing their own...
Boots lightweight Linux VMs using Apple's Virtualization.framework with ephemeral rootfs. Agents get a disposable environment to execute code without touching your host. Checkpoints, network access, port forwarding. If you run Claude Code on macOS and want proper isolation wit...
AMAP-ML/LongHorizon-Harness (370 stars, MIT, paper at arXiv 2608.01964) splits long computer-use work into three roles with independently assignable models. Reported: WeaveBench 51.8% → 80.7%, Terminal-Bench 2.1 69.7% → 77.2%, OSWorld 2.0 2.8% → 8.3%. It integrates with Claude...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.