Dispatch
AWS measures query-aware RAG compression at 8.6x fewer tokens for a 2.5-point quality drop
An August 21 AWS post benchmarks a two-call pattern on Bedrock: Claude Haiku extracts verbatim query-relevant spans from retrieved chunks at temperature 0.0, then Claude Sonnet answers from the filtered evidence. Compression alone sends 8.6x fewer tokens (12% of baseline) for 33% cost savings at 97.5% of baseline composite quality and 19% higher latency; adding a reranker gets to 10.1x fewer tokens and 36% savings at 12% higher latency. Hallucination rate dropped 7 points with compression and 13 points with rerank plus compression, which is the number worth stealing.
↳ Follow the thread