VICO:用于视觉语言模型推理的视觉环境协同进化框架
VICO: Visual Environments Co-Evolving for Vision-Language Model Reasoning
浏览论文内容
中文总结 AI 辅助
VICO是让视觉语言模型的智能体与环境重写器协同进化的框架,无需额外人工标注,在9个多模态基准测试中提升了域外任务性能,为视觉推理提供了可扩展路径。
中文摘要 AI 辅助
带有可验证奖励的强化学习(RLVR)已成为视觉语言模型(VLM)后训练的标准方法,但它通常假设训练环境是静态的。随着智能体(actor)性能提升,固定任务会偏离其学习前沿:许多任务变得过于简单,另一些则仍无法解决,学习信号会崩溃。我们认为VLM后训练应让视觉环境与智能体协同进化,而非仅优化智能体本身。我们提出VICO,这是一个协同进化框架,其中智能体与环境重写器(EnvRewriter)联合训练:EnvRewriter编辑可验证的图像侧结构,如图场景图、图表表格或保护区域掩码,并重新渲染生成标签有效的训练样本,其难度通过基于通过率的奖励校准至智能体当前能力。该循环无需额外人工标注,持续使任务难度与智能体能力重新对齐。在涵盖数学推理和视觉基础理解的9个多模态基准测试中,VICO-8B在域外任务上较其基础模型提升最高达5.0%,分别比最强的自进化基线和文本编辑协同进化基线高出4.3%和8.4%,且与使用16至160倍更少标注样本的图表专用RLVR方法表现相当。通过从人工标注监督转向图像编辑协同进化,VICO为超越静态语料库RLVR的视觉推理提供了可扩展路径。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) has become a standard recipe for post-training vision-language models (VLMs), but it typically assumes a static training environment. As the actor improves, fixed tasks drift out of its learning frontier: many become trivial, others remain unsolvable; and the learning signal collapses. We argue that VLM post-training should evolve the visual environment alongside the actor, not just the actor itself. We propose VICO, a co-evolutionary framework in which an actor and an Environment-as-Rewriter (EnvRewriter) are trained jointly: the EnvRewriter edits verifiable image-side structures, such as scene graphs, chart tables, or protected region masks, and re-renders them to produce label-valid training samples whose difficulty is calibrated to the actor's current ability through a pass-rate-based reward. This loop continuously realigns task difficulty with actor capability without any additional human annotation. Across nine multimodal benchmarks spanning mathematical reasoning and visually grounded understanding, VICO-8B improves over its base model by up to +5.0% on out-of-domain tasks, surpasses the strongest self-evolution and text-editing co-evolution baselines by +4.3% and +8.4% respectively, and stays comparable to chart-specialized RLVR methods using 16-160 times fewer labeled samples. By shifting from human-labeled supervision to image-editing co-evolution, VICO offers a scalable path beyond static-corpus RLVR for visual reasoning.
发表机构
- Virginia Tech(弗吉尼亚理工大学)
- Cisco(思科公司)
- NVIDIA(英伟达公司)
- UChicago(芝加哥大学)
- Georgia Tech(佐治亚理工学院)
- Abaka AI(Abaka人工智能公司)
- UC Berkeley(加州大学伯克利分校)
- UGA(佐治亚大学)
- WashU(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。