Reveals that standard response-level DPO for multimodal reasoning models performs similarly to answer-only optimization — the Chain-of-Thought supervision is insufficiently exploited. Proposes reasoning-conditioned preference optimization that separately targets CoT-level and answer-level preferences. From Tsinghua/HIT. Addresses a key failure mode as reasoning-augmented multimodal models become standard.