MCPO:用于多模态思维链压缩的模态对比偏好优化
MCPO: Modality-Contrastive Preference Optimization for Multimodal Chain-of-Thought Compression
浏览论文内容
中文总结 AI 辅助
针对多模态思维链过长导致的计算开销与KV缓存压力问题,提出MCPO两阶段压缩方法,在主流模型上实现最高69.5%的思维链长度缩减与3.34倍推理加速且保持准确率。
中文摘要 AI 辅助
近年来,多模态大规模推理模型通过长思维链(M-CoT)展现出解决复杂任务的卓越能力。然而,过长的推理轨迹会产生大量计算开销和显著的KV缓存压力。现有思维链压缩与对齐范式主要依赖静态规则或单维度偏好,缺乏细粒度的跨模态约束,因此容易引发视觉懒惰和幻觉推理。为解决这些问题,我们提出模态对比偏好优化(MCPO),这是一种样本效率极高的两阶段长度压缩方法,仅需不到900个训练样本。在压缩阶段,我们引入了步级归一化跨模态互信息(NCMI)剪枝算法,通过对比带图像与无图像上下文之间的推理差异,自动识别并移除与视觉无关的推理步骤,这显著减少了推理链中的冗余和幻觉内容。在对齐阶段,模型首先通过监督微调实现领域自适应初始化,随后采用非对称多模态长度控制偏好损失进行优化。该目标采用高度非线性的优势比公式,在带图像上下文提供陡峭梯度以强化首选轨迹的长度约束,同时在无图像上下文应用缩放后的平缓梯度线性差异以维持模态一致性,从而实现稳定的跨模态偏好对齐。在Qwen3-VL-Thinking等主流基础模型上开展的大量实验表明,我们的方法可将思维链长度最多减少69.5%,实现最高3.34倍的端到端推理加速,同时保持原始准确率。
英文摘要
Recently, multimodal large-scale reasoning models have demonstrated remarkable capabilities in solving complex tasks through long Chains-of-Thought (M-CoT). However, excessively long reasoning trajectories incur substantial computational costs and significant KV-cache pressure. Existing CoT compression and alignment paradigms mainly rely on static rules or single-dimensional preferences, lacking fine-grained cross-modal constraints; as a result, they are prone to inducing visual laziness and hallucinatory reasoning. To address these issues, we propose Modality-Contrastive Preference Optimization (MCPO), a highly sample-efficient two-stage length-compression method that requires fewer than 900 training samples. In the compression stage, we introduce a step-level Normalized Cross-Modal Mutual Information (NCMI) pruning algorithm, which automatically identifies and removes visual-independent reasoning steps by comparing the reasoning discrepancies between with-image and no-image contexts. This significantly reduces redundancy and hallucinatory content in the reasoning chains. In the alignment stage, the model first undergoes supervised fine-tuning to achieve domain-adaptive initialization, followed by optimization using an asymmetric multimodal length-controlled preference loss. This objective adopts a highly nonlinear odds-ratio formulation that provides steep gradients in the with-image context to reinforce length constraints for preferred trajectories, while applying a scaled, flat-gradient linear difference in the no-image context to maintain modality consistency, thereby achieving stable cross-modal preference alignment. Extensive experiments on mainstream base models such as Qwen3-VL-Thinking show that our method can reduce CoT length by up to 69.5% and achieve up to 3.34x end-to-end inference speedup while preserving original accuracy.
发表机构
- Tsinghua University(清华大学)
- Huawei Technologies Ltd.(华为技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。