arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.13854cs.CL

SPyCE:多模态智能体的技能-策略协同进化

SPyCE: Skill-Policy Co-evolution for Multimodal Agents

Ru Zhang, Weijie Qiu

首次发表
浏览论文内容

中文总结 AI 辅助

研究多模态智能体,提出SPyCE框架,将推理轨迹提炼为分层技能库与策略协同进化,执行技能捕获局部操作,工作流技能编码高级先验,实验证明该框架优于基线,为构建多模态智能体提供新范式。

中文摘要 AI 辅助

多模态智能体通过图像思考,在多个步骤中迭代地操作视觉证据并调用工具。现有强化学习方法将轨迹简化为标量奖励,迫使策略在每个新任务上从头发现可重复使用的工具使用模式;基于记忆的方法保留过去经验,但依赖测试时检索,未更新策略以吸收经验中的可重复模式。我们的关键见解是,多模态推理轨迹应提炼为在训练期间与策略协同进化的可重复使用技能,而非作为奖励消耗或从静态存储中检索。为此,我们提出SPyCE框架,将轨迹提炼为分层技能库并在强化学习中更新。执行技能捕获局部视觉操作,工作流技能编码协调工具使用的高级先验。训练期间,策略模型基于检索到的技能指导展开,技能库利用策略生成的有价值展开进行进化。实验表明SPyCE优于基于强化学习和基于记忆的基线。结果表明联合技能-策略优化是构建强大多模态智能体的有前景范式。

英文摘要

Multimodal agents that think with images iteratively manipulate visual evidence and invoke tools across many steps. Existing reinforcement learning methods reduce trajectories to scalar rewards, forcing the policy to discover reusable tool-use patterns from scratch on every new task; memory-based alternatives retain past experience, yet they rely on test-time retrieval, without updating the policy to absorb reusable patterns from that experience. Our key insight is that multimodal reasoning trajectories should be distilled into reusable skills that co-evolve with the policy during training, rather than being consumed as rewards or retrieved from a static store. To this end, we propose SPyCE (Skill-Policy Co-evolution), a framework that distills trajectories into a hierarchical skill library and updates it throughout reinforcement learning. Execution skills capture local visual operations, while workflow skills encode high-level priors that orchestrate tool use. During training, the policy model conditions on retrieved skills to guide its rollouts, while the skill library evolves using valuable rollouts generated by the policy. This creates a closed loop in which improved policies yield better skills, and the evolving skill library, in turn, provides stronger priors for policy rollouts. Experiments across eight benchmarks demonstrate that SPyCE consistently outperforms both RL-based and memory-based baselines. Further analysis reveals that both the hierarchical skill design and the co-evolution mechanism are critical to our design. These results suggest joint skill-policy optimization as a promising paradigm for building capable multimodal agents.

发表机构

  • Zhejiang University(浙江大学)
  • Beijing University of Posts and Telecommunications(北京邮电大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑