arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16459cs.LGcs.AIcs.CV

OPD-Aha:多模态在线策略蒸馏中从语言动量到视觉反思

OPD-Aha: From Linguistic Momentum to Visual Reflection in Multimodal On-Policy Distillation

Chenhao Qiu, Dawei Li, Yechao Zhang, Lei Gong, Zhen Tan

首次发表
浏览论文内容

中文总结 AI 辅助

针对多模态在线策略蒸馏中师生共同幻觉问题,提出OPD-Aha方法,从教师视觉偏好重建目标,抑制错误延续并促使学生反思,显著提升细粒度感知与复杂推理性能。

中文摘要 AI 辅助

特权在线策略蒸馏通过允许教师利用丰富的、仅训练时可用的视觉证据来评估学生轨迹,从而提升多模态推理能力。两个模型在基于相同的学生生成前缀进行条件生成时,都会对这些轨迹进行评分。当学生在响应早期误解图像时,这种累积的错误推理最终会将教师从其视觉证据上拉开。教师和学生收敛于相同的幻觉,导致标准的跨模型监督在最需要纠正的地方恰好失效。我们发现,在这种误导性一致下,教师的视觉纠正偏好并未丢失。比较同一教师在给定真实图像和视觉空值时的预测,可以发现特权证据仍然推动模型走向正确的解释。我们提出了OPD-Aha,它直接从这种孤立的视觉偏好中重建蒸馏目标,而不是依赖脆弱的教师-学生差异。这个重建的目标会积极抑制与图像矛盾的延续。通过这一目标训练,学生学会自然地用诸如“wait”和“actually”这样的反思标记打断自己错误的推理。反思之后,后续生成较少依赖累积的错误文本,而更多依赖视觉证据。在生成过程中纠正这些轨迹从根本上改变了推理过程,在多样的细粒度感知和复杂多模态推理基准上产生了广泛而一致的改进。我们的代码和模型可在以下网址获取:https URL。

英文摘要

Privileged on-policy distillation improves multimodal reasoning by allowing a teacher to evaluate student trajectories using rich, training-only visual evidence. Both models score these trajectories while conditioning on the same student-generated prefix. When a student misinterprets an image early in a response, this accumulating erroneous rationale eventually pulls the teacher away from its visual evidence. The teacher and student converge on the same hallucination, causing standard cross-model supervision to collapse precisely where correction is most needed. We find that the teacher's visual corrective preference is not lost under this misleading agreement. Comparing the predictions of the identical teacher given the real image and a visual null reveals that the privileged evidence still pushes the model toward the correct interpretation. We introduce OPD-Aha, which reconstructs the distillation target directly from this isolated visual preference rather than relying on the fragile teacher-student discrepancy. This reconstructed target aggressively suppresses continuations that contradict the image. Trained with this objective, students learn to naturally interrupt their own flawed reasoning with reflection tokens such as wait and actually. After reflection, subsequent generation relies less on the accumulated erroneous text and more on the visual evidence. Correcting these trajectories mid-generation fundamentally alters the reasoning process, yielding broad and consistent improvements across diverse fine-grained perception and complex multimodal reasoning benchmarks. Our code and models are available at https://github.com/Echochef/OPD-Aha.

发表机构

  • Arizona State University(亚利桑那州立大学)
  • University of Virginia(弗吉尼亚大学)
  • Stevens Institute of Technology(史蒂文斯理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑