arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Evidence-RL:面向证据密集型视觉推理

Evidence-RL: Towards Evidence-intensive Visual Reasoning

Haojie Huang, Xinlei Yu, Chengming Xu, Zhangquan Chen, Cheng Yang, Qingdong He, Yu Yang, Jiangning Zhang, Xiaobin Hu

arXiv 2608.08021首次发表:更新:

发表机构

National University of Singapore; Zhejiang University; Fudan University; Tsinghua University; Tencent(新加坡国立大学; 浙江大学; 复旦大学; 清华大学; 腾讯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视觉-语言模型依赖语言先验等问题,提出反事实证据解缠方法,结合GRPO框架在9个基准和4种骨干模型上优于现有基于RL的后训练方法。

AI 中文摘要

视觉-语言模型(VLMs)应基于图像中的具体证据进行回答,而非依赖语言先验、数据集捷径或无关视觉上下文。现有感知感知的后训练方法通过全局扰动或注意力代理来鼓励模型利用图像信息,但未测试采样的答案是否因果依赖于支持它的局部证据。我们提出反事实证据解缠(Counterfactual Evidence Disentanglement,CED),这是一种针对VLM grounding的训练时证据审计方法。对于每个回答,CED会中和以对象为中心的证据区域,并将由此产生的支持度下降与匹配的非证据区域进行比较。我们将该信号与GRPO框架内的答案正确性相结合,奖励那些依赖证据路径而非捷径或干扰路径的正确答案。CED使用弱对象级别的提议,无需特定问题的证据注释,且无推理时开销。在9个公共基准和4种骨干模型上,CED的表现优于现有的基于RL的后训练方法,针对性分析也验证了其以对象为中心的信号有效性。

英文摘要

Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visual context. Existing perception-aware post-training methods encourage image use through global perturbations or attention proxies, but they do not test whether a sampled answer causally depends on the local evidence that supports it. We propose Counterfactual Evidence Disentanglement (CED), a training-time evidence audit for VLM grounding. For each response, CED neutralizes an object-centric Evidence Region and compares the resulting support drop against matched non-evidence Regions. We combine this signal with answer correctness inside GRPO, rewarding correct answers that rely on the evidence path rather than shortcut or nuisance paths. CED uses weak object-level proposals, requires no question-specific evidence annotations, and adds no inference-time overhead. Across nine public benchmarks and four backbones, CED outperforms prior RL-based post-training methods, with targeted analyses verifying its object-centric signal.

Comments22 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑