Dispatch
AWS benchmarks show cheapest-per-token is the wrong metric: a pricier model cost 8x less per correct answer
AWS compared gpt-5.6-luna, terra and sol on Bedrock against gpt-5.4-mini and nano. On AIME, Luna cost $0.0021 per passing answer versus Mini at $0.0139 despite Mini's lower nominal token price. On DeepSearchQA multi-turn agent trajectories the gap widened to $0.05 vs $0.40 per passing answer with better quality (F1 0.50 vs 0.39), because every turn re-sends conversation context and cost accumulates roughly quadratically in turn count. On GDPval, Luna passed 56% at $0.010 per deliverable against Mini's 42% at $0.030.
↳ Follow the thread