Agents
Cognition ships SWE-2 on a Kimi K3 base: 92.8% on Terminal-Bench 2.1, and within a point of Fable 5.1 on FrontierCode at 64% lower cost
Cognition released SWE-2 on 2026-09-10. The model is RL-trained on the 2.8T-parameter Kimi K3 base, with medium, high and max effort levels trained in a single RL run. It scores 50.0% on FrontierCode 1.1 Main (Fable 5.1: 50.9%, GPT-6 Astra: 53.3%) and 73.0% on DeepSWE 1.1, and posts a leading 92.8% on Terminal-Bench 2.1. On the harder Terminal-Bench 4 it trails badly at 27.3%, against 55.8% for Fable 5.1 and 57.9% for GPT-6 Astra. It is available now in Devin Desktop, CLI, Web and Fusion. The Terminal-Bench 4 gap suggests the cost win holds on routine agent tasks and not on the longest-horizon ones.
Source
↳ Follow the thread