Fetching from the wire…
Public story · 2026-08-04 · high
The benchmark uses stateful replicas of real commerce services, not static transcripts, with reference scores run through the OpenClaw harness.
Why now: As of August 4, 2026, RealReplicaBench sits at v1.3.1, with a live leaderboard populated by OpenClaw reference results, per the repo.
Alibaba's Accio team published RealReplicaBench, 107 long-horizon agent tasks run in stateful replicas of real commerce services, per the project's GitHub repo. It matters most to people building agents, not people reading about them. Most agent benchmarks still score against static transcripts, which can't show whether an agent's earlier actions leave the world workable for what happens next.
RealReplicaBench targets that gap directly. Its replicas hold state across a task's full sequence, and reference results from the OpenClaw harness feed a live leaderboard instead of one static score. The tasks span long-horizon commerce workflows, not isolated single actions.
The project also ships a reproducibility contract alongside the 107 tasks, per the repo.
The repo sits at 162 stars on version 1.3.1, small by benchmark standards but real adoption for a tool this specific.
For builders shipping agents that touch real services, a static-transcript score has always been a soft proxy. It's never shown whether an agent can finish a multi-step job without breaking its own environment along the way.
Each link below shares sources, entities, or timing with this story.
Accio open-sourced it August 2, now 1,042 stars: 107 tasks across 14 containerized mock services replicating Alibaba, Shopify, and FreightOS (GitHub). Task mix is 53 CLI, 28 browser, 16 file ops, 10 API/MCP, split 65 text-only / 20 browser-text / 22 vision-requiring. Twelve mo...
Triple-stream retrieval (BM25 keyword, vector embeddings, knowledge-graph traversal) fused via Reciprocal Rank Fusion on the iii engine, with SQLite for state and an in-memory vector index, no external database. The economic claim: ~170K tokens/year (~$10) versus ~650K tokens...
The leaderboard says first place. The methodology says you should check your own bill. Qwen3.8 Max now ranks first on Artificial Analysis' agentic index, scoring 86.1 on OSWorld-Verified ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0, priced at $2.00/M input and $6.00/M...
A spec is a press release until someone who didn't write it implements it. GitHub made Agent Plugins 1.0 generally available on August 12 across VS Code, Copilot CLI, the Copilot SDK, and the Copilot app on all plans. The spec, published August 6, was co-authored by AWS, Anysp...
affaan-m/ECC (36.3k forks, MIT) bundles 67 agents, 284 skills, 94 legacy command shims, and "instincts", patterns learned from prior sessions with confidence scores that auto-recall when relevant, plus a .ecc/memory/ markdown vault that's explicitly cross-harness, so context s...
Qwen 3.5 is the first major model pretrained specifically for agentic multimodal workflows from the first training stage, not fine-tuned after the fact. 397B total / 17B active parameters (MoE architecture), 256K context window, 201 languages. Ships with Qwen Code (terminal ag...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.