arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用图像-问题依赖性改进VLM测试时强化学习

Harnessing Image Question Dependence for Better VLM Test-time Reinforcement Learning

Xinrui He, Ting-Wei Li, Junting Wang, Mengting Ai, Xinyu He, Hanghang Tong, Jingrui He

arXiv 2609.13296首次发表:更新:

发表机构

University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对VLM测试时强化学习中共识信号不可靠的问题,提出TTIQ框架,利用图像与问题的依赖性构建奖励,在多个VQA数据集上取得最优平均性能。

AI 中文摘要

测试时强化学习可以将视觉语言模型(VLM)适应到无标注的目标数据上,但其有效性从根本上受限于自生成学习信号的可靠性。为了评估基于共识的学习信号的可靠性,我们跨多个VQA数据集和模型规模分析了VLM测试时强化学习,揭示了两个局限性。首先,基于共识的测试时训练所带来的收益主要来自答案归一化而非内容修正。其次,许多初始VLM响应是错误的,原因是模型联合使用图像和问题的能力有限;从这些输出中得出的共识奖励可能保留由此产生的接地错误,而不是纠正它们。受此启发,我们提出了TTIQ,一个利用图像-问题依赖性来更好地适应VLM的测试时强化学习框架。TTIQ在原始图像-问题对及其图像消融和问题消融变体下对每个采样响应进行教师强制,利用由此产生的token级似然变化来估计对每个输入的依赖性。它将图像依赖性和问题依赖性与校准置信度相结合,构建一个响应级奖励,该奖励偏好联合接地的响应,并使用token级信号将更大的正向策略信用分配给由两个输入支持的token。这种设计偏好那些在图像和问题上联合接地且足够自信的响应,而不仅仅是流行的响应。在八个VQA数据集和多种VLM规模上的实验表明,TTIQ在每个模型规模上都取得了最佳平均性能。它还能跨VLM家族泛化,而在一个数据集上训练的模型无需进一步训练即可提高在未见数据集上的性能。

英文摘要

Test-time reinforcement learning can adapt vision-language models (VLMs) to unlabeled target data, but its effectiveness is fundamentally limited by the reliability of self-generated learning signals. To assess the reliability of consensus-based learning signals, we analyze VLM test-time reinforcement learning across diverse VQA datasets and model sizes, revealing two limitations. First, gains from consensus-based test-time training largely come from answer normalization rather than content correction. Second, many initial VLM responses are incorrect due to the model's limited ability to jointly use the image and the question; consensus rewards derived from these outputs may preserve the resulting grounding errors rather than correct them. Motivated by these, we propose TTIQ, a test-time reinforcement learning framework that harnesses image-question dependence for better vlm adaptation. TTIQ teacher-forces each sampled response under the original image-question pair and its image- and question-ablated variants, using the resulting token-level likelihood changes to estimate dependence on each input. It combines image and question dependence with calibrated confidence to construct a response-level reward that favors jointly grounded responses, and uses the token-level signals to assign greater positive policy credit to tokens supported by both inputs. This design favors responses that are jointly grounded in the image and the question and sufficiently confident, rather than merely popular. Experiments across eight VQA datasets and multiple VLM sizes show that TTIQ achieves the best average performance at every model scale. It further generalizes across VLM families, while models trained on one dataset improve performance on unseen datasets without further training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑