Reddit
GLM-5.3 Max Took Second on the Short Story Creative Writing Benchmark, and the Qualitative Report Names What Changed
Lechmazur's Short Story Creative Writing Benchmark, where models write to identical constrained briefs and LLM judges pick the stronger story in matched pairs, now places GLM-5.3 Max second. The release adds in-depth qualitative reports comparing six new models to their predecessors across 50 matched stories per pair. The stated difference between GLM-5.2 Max and 5.3 is structural rather than stylistic: 5.2 protagonists work alone in an agreeable world, 5.3 puts a second person in the room who withholds or judges. That is a concrete, testable prompt-level distinction, not a vibes score.
↳ Follow the thread