基于策略的视觉证据蒸馏
On-Policy Visual Evidence Distillation
浏览论文内容
中文总结 AI 辅助
针对视觉智能体中证据获取、读取和锚定错误传播问题,提出ReVuE方法,通过比较学生轨迹诊断首个失败并重加权token级蒸馏损失,在11个基准上优于现有OPD基线。
中文摘要 AI 辅助
视觉智能体通过将推理与图像操作交替进行来解决问题,而基于策略的蒸馏(OPD)从强大的教师模型为学生生成的交互轨迹提供指导。然而,图像操作会改变后续推理可用的证据,因此证据获取(Acquire)、读取(Read)或答案锚定(Ground)中的局部错误可能会在轨迹中传播并导致错误答案。现有的多模态OPD方法主要构建或对比原始图像的辅助视图以加强监督,但未显式建模学生动作、由此产生的观察结果与后续推理之间的联系。这限制了它们针对不同失败阶段提供定制化修正的能力。我们提出了视觉证据反思(ReVuE),一种用于视觉智能体的基于策略的蒸馏方法。ReVuE比较同一查询下多条学生生成的轨迹,总结观察到的视觉证据,并诊断Acquire、Read和Ground阶段中的第一个失败。由此产生的反思为教师模型提供了训练时的上下文。我们根据这些反思对教师预测的影响程度对token级蒸馏损失进行分组和重新加权。这种设计将轨迹级证据诊断转化为针对性的token级监督,引导学生改进其视觉证据获取和推理。在涵盖Qwen2.5-VL和InternVL3.5模型家族的11个基准测试中,ReVuE在感知、数学推理和通用任务的加权平均分数上优于所有评估的OPD基线。ReVuE还减少了推理和工具调用中的冗余,同时提高了工具调用准确性和任务准确性。代码可在以下网址获取:此https URL
英文摘要
Visual agents solve problems by interleaving reasoning with image operations, and on-policy distillation (OPD) provides guidance from a strong teacher on student-generated interaction trajectories. However, image operations change the evidence available for subsequent reasoning, so local errors in evidence acquisition (Acquire), reading (Read), or answer grounding (Ground) can propagate through the trajectory and lead to incorrect answers. Existing multimodal OPD methods primarily construct or contrast auxiliary views of the original image to strengthen supervision, without explicitly modeling the connections between student actions, resulting observations, and subsequent reasoning. This limits their ability to provide corrections tailored to different failure stages. We introduce Reflection on Visual Evidence (ReVuE), an on-policy distillation method for visual agents. ReVuE compares multiple student-generated trajectories for the same query, summarizes the observed visual evidence, and diagnoses the first failure across the Acquire, Read, and Ground stages. The resulting reflections provide training-time context for the teacher. We group and reweight token-level distillation losses according to how strongly these reflections affect the teacher's predictions. This design translates trajectory-level evidence diagnosis into targeted token-level supervision, guiding students to improve their visual evidence acquisition and reasoning. Across 11 benchmarks spanning the Qwen2.5-VL and InternVL3.5 model families, ReVuE outperforms all evaluated OPD baselines in weighted-average scores for perception, mathematical reasoning, and general tasks. ReVuE also reduces redundancy in reasoning and tool calls while improving tool-call accuracy and task accuracy. Code is available at https://github.com/sylvain-wei/ReVuE
发表机构
- Tencent(腾讯)
机构由 AI 辅助整理,请以论文原文为准。