Fetching from the wire…
Public story · 2026-07-14 · high
Claude Code solved 90 of 105 tasks, 85.71%, beating Codex on GPT-5.5 and JetBrains' Junie, which both scored 81.9%.
Why now: JetBrains published the benchmark on its Kotlin blog in July 2026.
JetBrains released an open benchmark of 105 real Kotlin repository tasks, and Claude Code solved 90 of them. Running Opus 4.7 at xhigh effort, it scored 85.71%, beating JetBrains' own Junie agent and Codex on GPT-5.5, both at 81.9%. For teams choosing a coding agent for JVM work, that gap is measured instead of assumed.
Each task requires reading an issue, finding the relevant code, and shipping a patch verified in a container test suite, per JetBrains. Junie ran at max effort, JetBrains' highest setting, and still trailed.
Most agent benchmarks run on Python, so Kotlin developers have had to guess whether the scores translate. This one doesn't blur the variables. It names the agent, the model, and the effort level as separate settings.
JetBrains doesn't say whether Claude Code's 85.71% holds if the task set grows past 105. Junie and Codex could close a four-point gap on the next run. That's the number to watch as more agents get tested against it.
Each link below shares sources, entities, or timing with this story.
Terminal-Bench 2.1 results (entries dated June 17) put Codex CLI on GPT-5.5 first at 83.4%, Claude Code on Fable 5 second at 83.1%, and Claude Code on Opus 4.8 at 78.9%. The asterisk matters more than the ranking: Fable 5 and Mythos 5 have been export-suspended since June 12,...
Three moves, two days, no coordination between them. August 10–11: GitHub shipped Ollama as a BYOK provider inside Copilot for JetBrains (GitHub Changelog). Unsloth released Unsloth Desktop with a command literally named unsloth start claude, which points Claude Code and Codex...
OpenAI launched the GPT-5.6 family on July 14: Sol (flagship), Terra (cost-optimized), and Luna (fast tier), live across ChatGPT, Codex, and the API the same day after a US-government-requested delay for security review. The numbers are loud. Sol scored 53.6 on Agents' Last Ex...
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
SkillSentry (arXiv 2608.09253) targets the gap where an agent completes a task under skill guidance then fails the same task on a repeat run. It defines a DSL for runtime guidance, initializes it from skill specs plus insights mined from historical successful *and failed* trac...
AI Now Institute researchers Boyan Milanov and Heidy Khlaaf demonstrated turning a coding agent doing vulnerability review into the execution vector, planting hidden binaries disguised as build artifacts alongside a deceptive README.md. The payload worked unchanged on Sonnet 5...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.