发表机构
University of Science and Technology of China; Zhongguancun Academy; Shandong University; Tsinghua University; Southeast University(中国科学技术大学; 中关村学院; 山东大学; 清华大学; 东南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VETTA通过共享轻量级评论家的独立头联合学习回合级和令牌级信用,结合优势与令牌残差进行PPO更新,在ALFWorld和WebShop上显著提升成功率并降低计算成本。
AI 中文摘要
多轮LLM智能体通常在多次交互中仅获得稀疏的任务反馈,同时逐令牌生成每个响应。这产生了两个相关的信用分配问题:哪些响应有助于达成结果,以及每个响应中哪些生成决策至关重要?现有方法通常只关注一个层面:回合级方法评估完整响应,但不区分其中的决策;令牌级方法可以跨回合传播反馈,但不显式建模每个响应的信用。这些互补的局限性促使在两个层面学习信用,并在单一策略更新中协调它们。我们引入了VETTA,一种信用分配方法,通过共享轻量级评论家上的独立头联合学习回合级和令牌级价值。VETTA沿两个时间序列计算优势,并将每个回合优势与响应内中心的令牌残差结合用于PPO更新。此外,为降低价值学习成本,评论家仅保留用于初始化演员的预训练检查点中的早期Transformer块。在两个具有挑战性的智能体基准ALFWorld和WebShop上,VETTA使用Qwen2.5-1.5B-Instruct相比PPO分别将成功率提高了37.5%和22.3%,使用Qwen2.5-7B-Instruct分别达到95.5%和76.0%的成功率。评论家深度比较进一步表明,在显著降低评论家侧计算量的同时保持强任务性能。这些结果表明,紧凑的共享评论家可以协调回合级和令牌级信用,以提高智能体性能,同时保持价值估计高效。代码可在该https URL获取。
英文摘要
Multi-turn LLM agents often receive sparse task feedback across several interactions, while generating each response token by token. This creates two related credit-assignment questions: which responses helped achieve the outcome, and which generation decisions mattered within each response? Existing methods typically focus on only one level: turn-level methods evaluate complete responses but do not distinguish the decisions within them; token-level methods can propagate feedback across turns but do not explicitly model credit for each response. These complementary limitations motivate learning credit at both levels and coordinating it in a single policy update. We introduce VETTA, a credit assignment method that jointly learns turn- and token-level values through separate heads on a shared lightweight critic. VETTA computes advantages along both temporal sequences and combines each turn advantage with a within-response-centered token residual for PPO updates. Furthermore, to reduce value-learning cost, the critic retains only early Transformer blocks from the pretrained checkpoint used to initialize the actor. On two challenging agent benchmarks, ALFWorld and WebShop, VETTA improves success rates over PPO by 37.5% and 22.3%, respectively, with Qwen2.5-1.5B-Instruct and achieves success rates of 95.5% and 76.0%, respectively, with Qwen2.5-7B-Instruct. Critic-depth comparisons further show strong task performance with substantially lower critic-side computation. These results suggest that a compact shared critic can coordinate turn- and token-level credit to improve agent performance while keeping value estimation efficient. Code is available at https://github.com/Jiaju-Chen/VETTA-official.
Comments16 pages, 6 figures