Fetching from the wire…
Public story · 2026-03-17 · source-backed
UK AI Safety Institute found model capability at 10M tokens jumped from 1.7 to 9.8 steps completed on a 32-step corporate network attack. Each 10x compute increase yields 59% more steps. No plateau found. Best run: 22/32 steps — 6 hours of a 14-hour human expert workload. Source
Each link below shares sources, entities, or timing with this story.
SWE-NFI builds 188 tasks from merged Python PRs and operationalizes non-functional improvement as 92 executable rules, cleanly separating "tests still pass" from "the code got better." Best agent: 70.0% functional correctness, 0.0-1.3 on structural improvement against a human...
ORCA-bench pairs a live OpenTelemetry-instrumented microservice system (six days of metrics, logs and traces via Prometheus, Jaeger and OpenSearch, plus full source access) with 1,079 RCA tasks varying report specificity and co-occurring faults. Best result across five frontie...
Thinkingbox is an MCP-compatible sandbox with isolated sessions, full execution traces, and outcome evaluation against terminal backend state, carrying 507 policy-conditioned workflows across retail, hospitality, auto insurance, neobank internal IT and consulting support (arXi...
VAKRA (arXiv 2608.12282) benchmarks agents against 8,000+ executable APIs across 62 domains, verifying by re-executing predicted calls against live endpoints. Accuracy falls to 50-51% on compositional APIs and degrades over 50% as depth grows. Failures concentrate in entity di...
The diagnosis in this paper is better than the fix, and the fix is very good. Recurrent memory agents fail at long context, but not for the reason most people assume. The bottleneck isn't capture. It's retention. Retention falls below 30% at 896K tokens because every consolida...
Best-in-class computer-use models scored 42% on OSWorld-Verified in early 2025. Today the leader (Claude Fable 5) scores 85%. The human tester baseline is roughly 72%. a16z published the aggregation on August 10, pulling from production interviews and llm-stats leaderboard dat...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.