Quesma's Terminal-Bench test found RTK made agent runs slightly more expensive, not 60-90% cheaper
Quesma·high signal
Quesma ran RTK on Terminal-Bench 2.1. Average cost per task went up 1% on Fable 5.0 and 17% on DeepSeek V4, and pass rates fell 1 to 2 points. Terminal output is only about 7% of Fable's input tokens and 26% of DeepSeek's. Extra agent turns cancelled the compression, and 44 of 58 DeepSeek tasks cost more. RTK's own `rtk gain` counts removed output bytes, not dollars. A July JetBrains A/B test also found RTK 7.6% more expensive at low effort, so two independent tests now contradict the 60-90% claim.