发表机构
Rochester Institute of Technology(罗彻斯特理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉语言模型缺乏合适训练数据的问题,提出Trace分类法引导环境,将任务构建分解,通过共享语义状态确定相关元素,在该环境上的RLVR提升了模型在外部基准测试中的表现,证明训练可迁移。
AI 中文摘要
具有可验证奖励的强化学习(RLVR)显著提升了语言模型推理能力,但其在视觉语言模型上的扩展受限于缺乏广泛、可精确验证和可重现的训练数据。我们引入了Trace,一种用于多领域视觉推理的分类法引导环境。Trace将任务构建分解为场景语法和可执行任务程序,分离视觉实现与答案计算。一个共享语义状态决定渲染图像、提示、类型化答案、验证器状态和可重放实例跟踪。由此产生的环境包含277种场景语法和11个视觉领域的1000个任务,具有可控的语义和视觉变化。在64000个Trace实例上进行RLVR,Qwen2.5-VL-3B在24个外部基准测试中的宏平均提高了3.51个百分点,Qwen2.5-VL-7B提高了4.06个百分点,证明广泛的程序训练可以超越生成的任务分布进行迁移。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) has substantially improved language-model reasoning, yet its extension to vision-language models remains constrained by the lack of training data that are simultaneously broad, exactly verifiable, and reproducible. We introduce Trace, a taxonomy-guided environment for multidomain visual reasoning. Trace factorizes task construction into a scene grammar and an executable task program, separating visual realization from answer computation. A shared semantic state determines the rendered image, prompt, typed answer, verifier state, and replayable instance trace. The resulting environment comprises 1,000 tasks over 277 scene grammars and 11 visual domains, with controlled semantic and visual variation. RLVR on 64,000 Trace instances improves the macro-average across 24 external benchmarks by 3.51 percentage points for Qwen2.5-VL-3B and 4.06 points for Qwen2.5-VL-7B, providing evidence that broad procedural training can transfer beyond the generated task distributions. Project page: https://maveryn.github.io/trace/.