Fetching from the wire…
Public story · 2026-08-23 · high
aaabench hands agents a real game engine and professional-grade conditions, then publishes the test setup but no scores.
Why now: Covered in the August 23 briefing alongside GamePhanes, a second long-horizon game-building eval.
aaabench gives coding agents a real game engine, Unreal, and professional conditions and time to build an open-world game. The repository, from developer ukanwat, has picked up 375 stars and 70 forks since July 31, per GitHub. What's missing is the point: no results, no leaderboard, no scored runs. Just the apparatus.
That's an unusual call for a benchmark. Most eval releases lead with a table of numbers because that's what gets cited. Withholding results while shipping the harness first is defensible here specifically because a benchmark this expensive to run needs the test validated before anyone trusts the scores it produces. Building an open-world game in a professional engine isn't a five-minute eval loop.
aaabench isn't alone. GamePhanes, a Godot-based agent environment and benchmark, hit 106 stars in two days, per its GitHub repository. Two long-horizon game-building evals surfacing side by side says something about where benchmark designers think agents are falling short: not on toy tasks, but on sustained, multi-system work.
Games make sense as the test bed because they split a question that most coding benchmarks collapse into one. Does the code run is one thing. Is the game any good, meaning is it playable, is it fun, does it hold together as a system, is a separate and harder thing. An agent can pass the first test and fail the second, and most benchmarks never find out because they stop checking after the code compiles.
What aaabench doesn't say yet is how it's grading "any good" once it does publish scores, or whether that judgment comes from a person, another model, or some fixed rubric. That's the detail that will decide whether this benchmark tells builders anything they didn't already assume.
Each link below shares sources, entities, or timing with this story.
Its loop is inspect, edit, run, observe, diagnose, repair, verify, and the evaluator drives controlled input probes against the running project to confirm runtime state changed, rather than stopping at files and exit codes. 224 stars since August 21, with one complete task pub...
Three frontier models shipped in a single week this month, and teams with a standing eval harness had a routing decision in hours. Anthropic's own agent-eval guidance says 20-50 tasks drawn from your real usage and real failures is enough to detect issues (DeepEval). DeepEval...
genspark-ai/genoffice (142 stars, created July 31, TypeScript, Apache-2.0) is a native macOS and Windows suite with word processor, spreadsheet, presentations and PDF, positioned as AI-native rather than AI bolted onto an existing editor. It lands alongside OfficeCLI (24,614 s...
yc-software/qm launched July 29 under MIT and is at 5,917 stars, 615 forks, 74 open issues: near 1,972 stars/day. Headless TypeScript/Fastify core on Postgres giving every person in an org an isolated sandbox with scoped memory, files and permissions, sharing work through Slac...
nocobase (~23,300 stars) explicitly rejects generate-everything-from-scratch, positioning AI as an operator on top of production-proven infrastructure. It's a direct rebuttal to cloudflare/vibesdk and datawhalechina/easy-vibe trending beside it, and it lands the same week Godo...
On July 1, the Godot Foundation rewrote its contributor guidelines to ban nearly all generative-AI code submissions and autonomous-agent PRs. New contributors with three or fewer merged PRs now need maintainer permission before submitting features or refactors. (The Register)...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.