LLMs that score strongly on isolated single-turn reasoning benchmarks show measurable degradation when the same tasks appear embedded in multi-turn dialogue contexts. The paper documents this effect empirically across multiple model families and task types, with conversational framing consistently suppressing reasoning accuracy. The direct implication: single-turn benchmark scores systematically overestimate real-world LLM reasoning performance in deployed conversational agents.