发表机构
University of Sydney; City University of Hong Kong(悉尼大学; 香港城市大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对MLLM细粒度感知,提出区域级策略优化方法Vision-RL2,通过强化学习优化提议网络,在减少视觉令牌的同时提升多基准准确性。
AI 中文摘要
多模态大语言模型(MLLMs)中的细粒度视觉感知通常通过提高分辨率来改进,但增加的视觉令牌会膨胀视觉编码和语言模型预填充成本。我们表明,细粒度感知所依赖的两个操作——定位感兴趣区域(RoI)和识别其内容——具有不同的分辨率要求。在受控诊断中,定位对令牌压缩的容忍度约为识别的3到4倍,这促使我们从粗略视图进行定位,并将分辨率集中在选定的证据上。使用MLLM解码坐标可以从答案端到端训练,但每次查询需要完整的模型传递,且依赖于接地能力。从模型注意力中蒸馏出的轻量级提议网络速度快,但继承了其注意力目标的噪声。来自提议网络的RoI通过离散的区域选择到达答案,因此其对答案的忠实度无法监督该网络。因此,我们使用区域级强化学习优化提议网络,称之为Vision-RL2。它将连贯区域视为动作,冻结的MLLM阅读器通过移除该区域后答案似然的变化来评分每个区域。互补的减法和加法目标抑制分散注意力的提议并恢复缺失的证据,仅更新预测器,无需区域标注、响应采样或推理轨迹。改进的提议进一步实现了稀疏编码,放大证据并排除背景令牌。在六个细粒度基准和四个MLLM骨干网络上,Vision-RL2在每个令牌预算下都优于基础模型的准确性,并以约4倍更少的视觉令牌超越了其最大预算的准确性。代码可在https://this URL获取。
英文摘要
Fine-grained visual perception in MLLMs is commonly improved by raising the resolution, but the added visual tokens inflate vision-encoding and language-model prefilling costs. We show that the two operations underlying fine-grained perception, localizing the region of interest (RoI) and recognizing its content, have different resolution requirements. In a controlled diagnostic, localization tolerates roughly 3 to 4 times stronger token compression than recognition, which motivates localizing from a coarse view and concentrating resolution on the selected evidence. Decoding coordinates with the MLLM can be trained end-to-end from answers, but costs a full model pass per query and depends on grounding ability. A lightweight proposal network distilled from the model's attention is fast, but inherits the noise of its attention targets. The RoI from the proposal network reaches the answer through a discrete region choice, so its faithfulness to the answer cannot supervise the network. We therefore optimize the proposal network with region-level reinforcement learning, which we call Vision-RL2. It treats coherent regions as actions, and a frozen MLLM reader scores each one by how its removal changes the answer likelihood. Complementary subtractive and additive objectives suppress distracting proposals and recover missing evidence, updating only the predictor without region annotations, response sampling, or reasoning trajectories. The refined proposal further enables a sparse encoding that magnifies evidence and excludes background tokens. Across six fine-grained benchmarks and four MLLM backbones, Vision-RL2 improves accuracy over the base model at every token budget and surpasses its largest-budget accuracy with about 4 times fewer visual tokens. Code is available at https://github.com/YuHengsss/VisionRL2 .