用于视觉语言模型智能体强化学习的统一评论家混合优势估计
Hybrid Advantage Estimation with Unified Critic for VLM Agentic Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
研究视觉语言模型智能体强化学习,提出HyGAE框架,用统一评论家估计值,推导混合优势联合优化令牌和轮次级目标,在多轮决策环境评估中平均成功率91%,比其他方法显著提升10%,证明混合优势解析形式对优化关键。
中文摘要 AI 辅助
大型视觉语言模型(VLM)如今在交互式环境中充当智能体,成功需要跨轮次的连贯推理和决策。虽然在智能体环境中的端到端训练可提升多轮决策能力,但当前方法主要依赖于对连接的令牌轨迹进行逐令牌优化或对轮次内均匀信用进行逐轮优化。本文为这两种优化水平建立理论公式,推导了兼顾两个目标的混合优势。通过适当选择折扣因子和学习目标,证明统一评论家模型可估计逐轮和逐令牌的值。进而提出HyGAE,一个用混合优势和统一评论家联合优化令牌和轮次级目标的演员评论家框架。在五个多轮决策环境中对HyGAE进行广泛评估,其平均成功率达91%,比其他方法显著提高10%。还深入分析表明混合优势和回报的精确解析形式对优化至关重要。
英文摘要
Large Vision-Language Models (VLMs) now act as agents in interactive environments, where success requires coherent reasoning and decision-making across turns. Although end-to-end training in agentic environments can improve such multi-turn decision-making abilities, current methods mainly rely on either token-wise optimization over concatenated token trajectories or turn-wise optimization with uniform within-turn credit. In this work, we establish theoretical formulations for the two levels of optimization and derive a hybrid advantage that serves both objectives. Furthermore, with an appropriate choice of discount factor and learning target, we prove that a unified critic model can estimate values for both turn-wise and token-wise. As such, we propose HyGAE, an actor-critic framework that jointly optimizes token- and turn-level objectives with the hybrid advantage and unified critic. We conduct extensive evaluations of HyGAE across five multi-turn decision-making environments, where it achieves an average success rate of 91% and a significant improvement of 10% over other methods. Furthermore, we provide an in-depth analysis showing that the exact analytic form of the hybrid advantage and return is crucial for optimization. Project Page: https://wx-zhang.github.io/hygae-web/.
发表机构
- KAUST(沙特阿卜杜拉国王科技大学)
机构由 AI 辅助整理,请以论文原文为准。