发表机构
University of Chinese Academy of Sciences(中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
V-Gym提出技能-数据协同进化框架,从执行轨迹中迭代提炼技能并生成针对性练习数据,在多个多模态推理基准上显著超越基线,实现自主诊断与持续自我提升。
AI 中文摘要
多模态理解、推理和工具使用方面的进展使智能体能够处理日益复杂的视觉推理任务。通过将过去的执行经验提炼为可复用的技能,智能体可以将成功和失败中的教训转移到未来的推理中,减少重复错误并提升能力。然而,有限的经验可能产生不可靠、泛化能力差的技能,而静态数据集可能缺乏针对性且多样化的练习来促进技能改进。为解决这一差距,我们提出了V-Gym,一个自主框架,它从执行轨迹中迭代地协同进化程序性技能和多模态练习数据。在技能进化过程中,V-Gym分析轨迹以提炼和优化层级技能,更新程序性指导和适用条件,并且仅在提升验证性能时才保留更新。在数据进化过程中,V-Gym通过平衡数据效用和探索来选择生成种子,然后将轨迹识别出的瓶颈转化为多样化、有针对性的练习数据,这些数据在质量检查后扩充数据银行。由此产生的练习结果反馈到后续的技能更新中,形成持续技能改进的闭环。在多个多模态推理基准上的实验表明,与多个骨干模型的基线相比,V-Gym取得了显著改进。其进化出的技能可跨领域和模型泛化,而进化出的数据支持更有效的技能改进,从而实现自主诊断、针对性练习和持续自我提升。
英文摘要
Advances in multimodal understanding, reasoning, and tool use enable agents to tackle increasingly complex visual reasoning tasks. By distilling past execution experience into reusable skills, agents can transfer lessons from both successes and failures into future reasoning, reducing repeated errors and improving capabilities. However, limited experience may produce unreliable, poorly generalizable skills, while static datasets may lack the targeted and diverse practice needed for refinement. To address this gap, we introduce V-Gym, an autonomous framework that iteratively co-evolves procedural skills and multimodal practice data from execution trajectories. During skill evolution, V-Gym analyzes trajectories to distill and refine hierarchical skills, updating procedural guidance and applicability conditions while retaining an update only if it improves validation performance. During data evolution, V-Gym selects generation seeds by balancing data utility and exploration, then translates trajectory-identified bottlenecks into diverse, targeted practice data that expand the data bank after quality checks. The resulting practice outcomes feed back into subsequent skill updates, closing the loop for continual skill refinement. Experiments across diverse multimodal reasoning benchmarks show substantial improvements over baselines with multiple backbone models. Its evolved skills generalize across domains and models, while evolved data support more effective skill refinement, enabling autonomous diagnosis, targeted practice, and continual self-improvement.