Fetching from the wire…
Top 5 · 2026-08-11 · source-backed
Best-in-class computer-use models scored 42% on OSWorld-Verified in early 2025. Today the leader (Claude Fable 5) scores 85%. The human tester baseline is roughly 72%. a16z published the aggregation on August 10, pulling from production interviews and llm-stats leaderboard data.
Crossing the human baseline is the eye-catching part, but the cost math is what changes decisions. Pure screenshot-loop operation runs $6–8/hour of inference, with the full range spanning $3–15 depending on how the harness is designed. Offshore BPO runs about $10/hour fully loaded. US back-office labor runs $30–45/hour. So the agent is roughly break-even against offshore and carries a 70–80% gross margin against domestic.
That's a very specific place to be. Not "cheaper than everything," which would have triggered instant commoditization. Not "still too expensive," which would have kept this in demo-land. Break-even against the cheapest human option and profitable against the expensive one, right now, at today's prices, which are falling.
The strategic read from the piece is the one builders should internalize: raw UI navigation has commoditized down into the model layer. If your product's value proposition was "we can click through a legacy web app reliably," the model does that now for $7/hour. The advantage moved up the stack to context (knowing which app, which account, which state), permissions (what the agent may touch), process knowledge (what the workflow actually is versus what the SOP says), validation (did it work), escalation (when to get a human), and run-caching (don't pay twice for the same trajectory).
One founder quote in the piece deserves attention for its precision: "the models weren't good enough to use in production on their own until Opus 4.6 in February 2026." That's six months ago. An entire product category became viable half a year ago, which means most of the durable companies in it haven't been founded yet.
This converges with Ouroboros reporting 90.69% on OSWorld-Verified from a completely different direction. Two independent sources putting computer-use well past the human baseline in the same week is the kind of agreement that's hard to dismiss as leaderboard gaming.
What I'd do with this: stop building the clicking. Start building the wrapper. If you have a workflow automation product, your roadmap for the next two quarters is process capture, permission scoping, and failure escalation, not better selectors. And if you're evaluating whether to buy or build here, note that $6–8/hour is a real operating cost that scales linearly, not a fixed-cost software play. The unit economics look more like staffing than like SaaS.
Each link below shares sources, entities, or timing with this story.
For two years the technique was accumulation. Longer system prompts, longer CLAUDE.md, more numbered do/don't lists, more "always verify your work" imperatives. Anthropic's context-engineering guidance for Claude 5 models inverts it, with an 80% deletion figure attached. The s...
Anthropic launched Claude Sonnet 4.6 claiming performance comparable to Opus 4.5 at $3/$15 per million tokens (vs. Opus's $5/$25). SWE-bench Verified: 79.6% (near Opus 4.6's 80.8%). OSWorld-Verified: 72.5% (tied with Opus 4.6's 72.7%). 1M token context window in beta. Now the...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
This one rearranged my week. An essay published August 4 walks through Databricks' independent benchmark of coding harnesses against its own multi-million-line codebase. Pi, a harness with four built-in tools and a system prompt under 1,000 tokens, paired with Opus 4.8 at xhig...
The core team released lemans on August 24 after deciding the Ruby community shouldn't have to run Python-based Harbor, and benchmarked four models on 63 Rails tasks (Rails). ox-alpha 52/63; Terra 49/63 at $0.20 and a 182-second median; open-weight Qwen 3.8-27B 48/63 but at a...
This is the paper of the week. arXiv 2607.28871 introduces BSG-VA, which replays every validation command an agent runs across three code states: the original buggy code (B), the candidate patch (S), and the gold developer fix (G). If a test passes in all three states, it neve...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.