When LLMs Stop Following Steps: Procedural Execution Fails Even When Final Answers Are Correct
arXiv·high signal
Researchers from arXiv introduce a diagnostic framework showing that LLMs often produce correct final answers while silently skipping or reordering intermediate procedural steps. The study demonstrates that final-answer accuracy alone is an unreliable proxy for whether models actually follow multi-step instructions — a critical finding for anyone building agentic workflows where step order matters (tool chains, pipelines, SOPs).