arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34557cs.AI

SkillRubric:多模态智能体的行动者指导与评估者准则的协同演化

SkillRubric: Co-Evolving Actor Guidance and Evaluator Rubrics for Multimodal Agents

Bingqing Jiang, Guoxi Zhang, Jasper Wang, Auric Wang, Bingning Wang, Tianyi Lin, Zichao Yu, Yujin Han, Ziye Ma, Difan Zou

首次发表
浏览论文内容

中文总结 AI 辅助

SkillRubric通过将技能表示为对齐的行动者指导和评估者准则,并采用交替协同演化方案,在多模态智能体基准上持续提升规划和工具使用性能。

中文摘要 AI 辅助

近期的工作将从过去交互中提炼的可复用技能融入多模态智能体训练,为长时程规划和工具使用提供程序性指导。然而,这些方法中的策略优化仍主要受稀疏的结果奖励驱动,对中间决策提供的监督很少。基于准则的奖励通过显式的中间标准解决了这一局限,但可靠的准则难以大规模构建,且往往与行动者所遵循的程序脱节。我们观察到,结构良好的技能自然规定了如何行动以及成功执行应达到什么效果。基于这一见解,我们提出了SkillRubric,它通过对齐的行动者导向指导和评估者准则来表示每个技能。一个多模态验证器利用截图和工具输出评估技能定义的目标,为负责的回合分配完成奖励和进度奖励。我们进一步提出了一种交替协同演化方案,该方案在冻结策略下通过配对回放验证指导修订,并在固定指导下离线验证准则修订。在多个多模态智能体基准上的实验表明,性能持续提升,而受控的配对回放进一步表明,演化后的技能比其前身版本为规划和工具使用提供了更有效的指导。

英文摘要

Recent work incorporates reusable skills distilled from past interactions into multimodal agent training, providing procedural guidance for long-horizon planning and tool use. However, policy optimization in these methods remains driven primarily by sparse outcome rewards, providing little supervision for intermediate decisions. Rubric-based rewards address this limitation through explicit intermediate criteria, but reliable rubrics are difficult to construct at scale and often disconnected from the procedure followed by the actor. We observe that a well-structured skill naturally specifies both how to act and what successful execution should achieve. Based on this insight, we introduce SkillRubric, which represents each skill through aligned actor-facing guidance and an evaluator-facing rubric. A multimodal verifier evaluates skill-defined goals using screenshots and tool outputs, assigning completion and progress rewards to the responsible turns. We further introduce an alternating co-evolution scheme that validates guidance revisions through paired rollouts under a frozen policy and rubric revisions offline under fixed guidance. Experiments across diverse multimodal agent benchmarks demonstrate consistent performance gains, while controlled paired rollouts further show that evolved skills provide more effective guidance for planning and tool use than their preceding versions.

发表机构

  • The University of Hong Kong(香港大学)
  • Peking University(北京大学)
  • WeChat, Tencent(腾讯微信)
  • The Hong Kong Polytechnic University(香港理工大学)
  • City University of Hong Kong(香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑