Fetching from the wire…
Public story · 2026-08-24 · high
BC-Bench evaluated 101 bug-fix tasks from two production Business Central repos written in AL, a DSL with little public training data.
Why now: The paper posted August 24, 2026.
Microsoft researchers built a 101-task bug-fixing benchmark from two production repos written in AL, the DSL behind Dynamics 365 Business Central. They adapted SWE-bench's method for a language with almost no public training data to learn from.
The results cut against a growing assumption that agent harness design matters more than model choice. In AL bug-fixing, gaps between frontier models exceeded gaps between the two harnesses tested, per the BC-Bench paper.
Gains models showed on general-purpose coding benchmarks didn't consistently carry over to AL bug-fixing either, the study found.
That's a problem beyond Business Central shops. Teams building on any DSL with thin public training data can't assume a model's benchmark rank holds for their codebase. For teams evaluating AI coding assistants, the difference matters. A model topping general leaderboards isn't guaranteed to be the best pick for a legacy or niche-language codebase.
The paper doesn't say why the harness effect shrank on AL specifically, only that it did in this test.
Each link below shares sources, entities, or timing with this story.
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
Three moves, two days, no coordination between them. August 10–11: GitHub shipped Ollama as a BYOK provider inside Copilot for JetBrains (GitHub Changelog). Unsloth released Unsloth Desktop with a command literally named unsloth start claude, which points Claude Code and Codex...
This is the other half of the Fable 5 story, so read them together. While the best coding model in the world is uncallable, an open-weight one quietly posted frontier-adjacent numbers. Per Tom's Hardware, independent benchmarks for the MIT-licensed GLM-5.2 (744B params, 40B ac...
MAI-Code-1-Flash, a 5B-parameter coding model, is in GitHub Copilot and VS Code, and Microsoft says it beats Claude Haiku 4.5 across core coding benchmarks, +16 points on SWE-Bench Pro at 51.2% versus 35.2%, using up to 60% fewer tokens. MAI-Thinking-1, a 35B-active MoE with a...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
Microsoft's July 23 release targets a genuine gap: harness-based agents like Claude Code and Codex drive multi-turn reasoning, tool use, and external system access but were hard to train end-to-end with standard open RL infrastructure. The trick is decoupling training from inf...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.