Fetching from the wire…
Research2026-06-18 · source-backed
Built with 250-plus industry experts, ALE evaluates agents on 960 expert-authored, deterministically-scored workflows across 55 industries. Agents average 26% across all tiers but 2.6% on the "last-exam" tier. The kicker: Codex with GPT-5.5 scores 82% on Terminal-Bench and 0% on Last-Exam tasks. The next time someone says agents are "job-ready," this is the number to quote back. (arXiv 2606.05405)
Each link below shares sources, entities, or timing with this story.
Terminal-Bench 2.1 results (entries dated June 17) put Codex CLI on GPT-5.5 first at 83.4%, Claude Code on Fable 5 second at 83.1%, and Claude Code on Opus 4.8 at 78.9%. The asterisk matters more than the ranking: Fable 5 and Mythos 5 have been export-suspended since June 12,...
OpenAI shipped GPT-5.5 on April 23, six weeks after 5.4. The capability jump is real: 82.7% on Terminal-Bench 2.0 vs Claude Opus 4.7's 69.4%. The Pro tier nearly doubles Opus 4.7 on FrontierMath Tier 4 at 39.6% vs 22.9%. It uses 40% fewer tokens on Codex tasks while matching 5...
SkillSentry (arXiv 2608.09253) targets the gap where an agent completes a task under skill guidance then fails the same task on a repeat run. It defines a DSL for runtime guidance, initializes it from skill specs plus insights mined from historical successful *and failed* trac...
OpenAI's changelog shows a Codex refresh several rankings now place at #1 for terminal-driven agentic coding. The release also optimizes TUI startup and session restore by querying the state DB first. The coding-agent tier keeps compressing. No single tool is safely ahead for...
OpenAI rolled GPT-5.5 into Codex for all paid tiers with 58.6% on SWE-Bench Pro. Codex now runs natively on Windows with PowerShell support, no WSL required. Pro subscribers get doubled usage through May 31.
Two thirds. Not two thirds of a contrived jailbreak set. Two thirds of realistic malicious issue requests, against the exact three tools most of the people reading this run daily. Ankur Singh, Jinqiu Yang, and Tse-Hsun Chen built IssueTrojanBench across four attack categories...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.