TIMETOACT's August Benchmark Compared 187 Models and Caught Claude Fable 5 Dropping From 90 to 83 on Re-Evaluation
TIMETOACT Group·low signal
The monthly TIMETOACT LLM benchmark round covers 187 models, and its most useful result is not the leaderboard but the re-evaluation finding: Claude Fable 5 scored 90 previously and 83 on retest, with no name or version change. That is direct evidence that a pinned model string is not a pinned capability, which matters for anyone whose evals or production prompts assume stability across a model alias. Single-source benchmark, so treat the absolute scores as indicative rather than authoritative.