arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10782cs.CVcs.AI

VICO:用于视觉语言模型推理的视觉环境协同进化框架

VICO: Visual Environments Co-Evolving for Vision-Language Model Reasoning

Meng Lu, Ligeng Zhu, Olivia Xiao, Yuchen Zhuang, Zihan Wang, Kuncheng Wu, Bangya Liu, Yu Wang, Charles Fleming, Wenqi Shi, Xuan Wang

首次发表
浏览论文内容

中文总结 AI 辅助

VICO是让视觉语言模型的智能体与环境重写器协同进化的框架,无需额外人工标注,在9个多模态基准测试中提升了域外任务性能,为视觉推理提供了可扩展路径。

中文摘要 AI 辅助

带有可验证奖励的强化学习(RLVR)已成为视觉语言模型(VLM)后训练的标准方法,但它通常假设训练环境是静态的。随着智能体(actor)性能提升,固定任务会偏离其学习前沿:许多任务变得过于简单,另一些则仍无法解决,学习信号会崩溃。我们认为VLM后训练应让视觉环境与智能体协同进化,而非仅优化智能体本身。我们提出VICO,这是一个协同进化框架,其中智能体与环境重写器(EnvRewriter)联合训练:EnvRewriter编辑可验证的图像侧结构,如图场景图、图表表格或保护区域掩码,并重新渲染生成标签有效的训练样本,其难度通过基于通过率的奖励校准至智能体当前能力。该循环无需额外人工标注,持续使任务难度与智能体能力重新对齐。在涵盖数学推理和视觉基础理解的9个多模态基准测试中,VICO-8B在域外任务上较其基础模型提升最高达5.0%,分别比最强的自进化基线和文本编辑协同进化基线高出4.3%和8.4%,且与使用16至160倍更少标注样本的图表专用RLVR方法表现相当。通过从人工标注监督转向图像编辑协同进化,VICO为超越静态语料库RLVR的视觉推理提供了可扩展路径。

英文摘要

Reinforcement learning with verifiable rewards (RLVR) has become a standard recipe for post-training vision-language models (VLMs), but it typically assumes a static training environment. As the actor improves, fixed tasks drift out of its learning frontier: many become trivial, others remain unsolvable; and the learning signal collapses. We argue that VLM post-training should evolve the visual environment alongside the actor, not just the actor itself. We propose VICO, a co-evolutionary framework in which an actor and an Environment-as-Rewriter (EnvRewriter) are trained jointly: the EnvRewriter edits verifiable image-side structures, such as scene graphs, chart tables, or protected region masks, and re-renders them to produce label-valid training samples whose difficulty is calibrated to the actor's current ability through a pass-rate-based reward. This loop continuously realigns task difficulty with actor capability without any additional human annotation. Across nine multimodal benchmarks spanning mathematical reasoning and visually grounded understanding, VICO-8B improves over its base model by up to +5.0% on out-of-domain tasks, surpasses the strongest self-evolution and text-editing co-evolution baselines by +4.3% and +8.4% respectively, and stays comparable to chart-specialized RLVR methods using 16-160 times fewer labeled samples. By shifting from human-labeled supervision to image-editing co-evolution, VICO offers a scalable path beyond static-corpus RLVR for visual reasoning.

发表机构

  • Virginia Tech(弗吉尼亚理工大学)
  • Cisco(思科公司)
  • NVIDIA(英伟达公司)
  • UChicago(芝加哥大学)
  • Georgia Tech(佐治亚理工学院)
  • Abaka AI(Abaka人工智能公司)
  • UC Berkeley(加州大学伯克利分校)
  • UGA(佐治亚大学)
  • WashU(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

↑