Fetching from the wire…
Top 5 · 2026-07-13 · source-backed
Here's the number that should sit in every "agents will replace engineers" thread: 15.2%. That's the best model. The mean across 15 frontier models is 4.3%.
Tencent Hunyuan's Long-Horizon-Terminal-Bench put 15 frontier models against 46 long-horizon terminal tasks across nine categories: experiment reproduction, software engineering, multimodal analysis, interactive games, scientific computing (arXiv:2607.08964, HuggingFace Daily Papers #1 today with 41 votes). It uses dense reward grading that scores partial progress through graded subtasks instead of pass/fail. The strongest model reached 15.2% pass@1 at a 0.95 partial-reward threshold, dropping to 10.9% at a perfect 1.0. Cross-model mean: 4.3%, falling to 1.7% at perfection.
And these tasks are brutally expensive. Roughly 9.9M tokens, 231 episodes, and 85 minutes of execution each. So the frontier isn't just failing at these. It's failing at them slowly and at enormous cost.
I love this paper because it's the counterweight to everything else in this issue. Grok's cheaper. Salesforce automated half its support. Tao builds apps with agents. And then this: the moment you point an agent at a genuinely long-horizon terminal task with real state and real dependencies, it falls off a cliff. The gap between "cheap and useful for scoped work" and "autonomous on hard multi-step problems" is not closing as fast as the pricing pages suggest.
This tracks with what I hit in my own projects. Agents are excellent at bounded tasks with clear success criteria. Give them something that requires holding state across 200 steps, recovering from a failed action, and reasoning about what went wrong three moves ago, and reliability collapses. The Mosaic paper this week found the same root cause: failed actions from inaccurate state tracking under partial observability are the dominant latency and failure bottleneck, not raw model speed (arXiv:2607.09603).
What builders should do: scope aggressively. Don't hand an agent a task that needs 200 coherent steps. Decompose it into bounded units with checkable outputs, and put the state in files the next run re-reads, not in the context window. The 15% number is your reminder that "autonomous" is a spectrum, and most of the useful zone is on the short-horizon end.
Each link below shares sources, entities, or timing with this story.
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill. SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-...
TechCrunch named the phenomenon everyone's been watching: the SaaSpocalypse. February saw $285B wiped from software stocks driven by three simultaneous forces — AI agents reducing headcount (fewer seats), coding agents making build-vs-buy favor build, and AI model providers mo...
HubSpot beat Q2 revenue at $912M, up 20%. Beat EPS at $3.26 against $3.02 expected. Its Data Agent is at 16,000 customers, up 80% quarter over quarter. Its Customer Agent resolves 72% of tickets without escalation. The stock had its largest one-day drop in 12 years on the NYSE...
This is the most consequential architecture decision in enterprise software since cloud versus on-prem, and it happened quietly across three vendor announcements. PYMNTS connected the dots first. SAP blocks. Its API Policy v4/2026, published in late April, prohibits using SAP...
Martin Alderson's essay "The upcoming AI margin collapse, part 1: GLM 5.2" hit 675 points and 462 comments on Hacker News, and it's the rare HN chart-topper that's actually about spreadsheet math instead of vibes. The argument is simple. Z.ai's GLM 5.2 delivers frontier-adjace...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.