TTRSD:视觉语言模型的测试时强化学习与自蒸馏
TTRSD: Test-Time Reinforcement Learning with Self-Distillation for Vision-Language Models
浏览论文内容
中文总结 AI 辅助
TTRSD通过多视角自蒸馏与视觉对比令牌选择,在测试时强化学习中将更新方向与位置分离,仅用20个无标签样本显著提升VLM性能,并保持推理完整性。
中文摘要 AI 辅助
测试时强化学习使视觉语言模型(VLM)能够利用无标签输入进行自适应。然而,在固定视觉条件下重复采样可能会强化共享的感知错误,而序列级奖励则无法隔离视觉感知——这一锚定多模态推理的基础瓶颈,从而有使预训练推理能力退化的风险。我们提出TTRSD,一种结合多视角答案级自蒸馏与视觉对比令牌选择的测试时强化学习框架。一个共享策略将教师模型在原始、裁剪和下采样视图上的预测聚合为答案分布。由原始图像生成的学生轨迹根据其最终答案在该分布中的支持度获得奖励。为了将该反馈精确地分配到感知瓶颈,我们在保持文本前缀固定的情况下,比较相同采样令牌在原始输入和视觉消融输入下的对数概率,选择视觉敏感位置进行策略梯度更新。TTRSD将更新方向(由组相对优势决定)与更新位置(由视觉敏感性决定)分离,无需真实标签、外部验证器或单独的教师模型。仅使用20个无标签自适应样本,TTRSD在七个基准和三个VLM上提升了性能,将InternVL3-2B的MMMU准确率从35.79%提高到49.32%(+13.53%),展示了跨数据集泛化能力,同时保持了固有的推理完整性。
英文摘要
Test-time reinforcement learning enables vision-language models (VLMs) to adapt using unlabeled inputs. However, repeated sampling under fixed visual conditions can reinforce shared perceptual errors, while sequence-level rewards fail to isolate visual perception the foundational bottleneck that anchors multimodal reasoning risking the degradation of pre-trained reasoning capabilities. We propose TTRSD, a test-time reinforcement learning framework combining multi-view answer-level self-distillation with visual contrastive token selection. A shared policy aggregates teacher predictions across original, cropped, and downsampled views into an answer distribution. Student trajectories generated from the original image receive rewards based on the support for their final answers in this distribution. To allocate this feedback precisely toward perceptual bottlenecks, we compare the log-probabilities of the same sampled tokens under original and visually ablated inputs while holding their textual prefixes fixed, selecting visually sensitive positions for policy-gradient updates. TTRSD separates update direction, determined by group-relative advantages, from update position, determined by visual sensitivity, without requiring ground-truth labels, external verifiers, or a separate teacher. With only 20 unlabeled adaptation samples, TTRSD improves performance across seven benchmarks and three VLMs, raising InternVL3-2B's MMMU accuracy from 35.79% to 49.32%(+13.53%), demonstrating cross-dataset generalization while preserving inherent reasoning integrity.
发表机构
- Zhejiang University(浙江大学)
- Hong Kong University of Science and Technology(香港科技大学)
- Harbin Institute of Technology(哈尔滨工业大学)
- Anhui University(安徽大学)
机构由 AI 辅助整理,请以论文原文为准。