arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VICT:用于长 horizon LLM 智能体强化学习的验证器插装式信用追踪

VICT: Verifier-Instrumented Credit Tracing for Long-Horizon LLM Agent Reinforcement Learning

Pengcheng Li, Zhengyang Zhang, Dongxu Zhang, Sui Huang, Shaohua Ma

arXiv 2608.28128首次发表:更新:

发表机构

Tsinghua University; Xi’an Jiaotong University; Jiaxing Nanhu University(清华大学; 西安交通大学; 嘉兴南湖学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长 horizon LLM 智能体强化学习的细粒度信用分配挑战,提出 VICT 方法,通过验证器侧的证明边追踪分配信用,在 ALFWorld 和 WebShop 上性能显著优于仅结果训练。

AI 中文摘要

细粒度信用分配是长 horizon LLM 智能体强化学习的核心挑战。标准目标通常基于可程序化验证的终端奖励训练,将每个稀疏结果广播到轨迹中的所有动作。现有方法通常从 rollout 侧寻求更精细的信用,构建辅助轨迹信号或额外比较来估计动作重要性。尽管这些方法有用,但仍将判断成功的验证器视为标量奖励,丢弃了其内部任务结构。我们的关键见解是,许多可验证任务已在其终端验证器中编码了相关检查。我们提出 VICT(Verifier-Instrumented Credit Tracing,验证器插装式信用追踪),这是一种训练时接口,可暴露可执行或证据支持的原子,并通过依赖验证的证明边将其追溯到动作。VICT 仅沿这些边重新分配组相对优势,将信用分配从 rollout 侧推理转移到验证器侧追踪。它保留原始终端奖励,在证据不完整或模糊时弃权(不执行),且仅改变训练时优势张量,无需学习的 critic、过程标签、分支 rollout 或推理时验证器访问。在 ALFWorld 和 WebShop 上,VICT 相比仅基于结果的训练有显著提升,且与近期细粒度信用方法相比实现了强劲性能;消融实验排除了密集原子奖励、最终提交信用、时间邻近性和稀疏性作为充分解释。

英文摘要

Fine-grained credit assignment is a central challenge in reinforcement learning for long horizon LLM agents. Standard objectives often train from programmatically verifiable terminal rewards by broadcasting each sparse outcome to every action in a trajectory. Existing methods typically seek finer credit from the rollout side, constructing auxiliary trajectory signals or additional comparisons to estimate action importance. Although useful, these approaches still treat the verifier that judged success as a scalar reward, discarding its internal task structure. Our key insight is that many verifiable tasks already encode the relevant checks inside their terminal verifier. We propose VICT (VerifierInstrumented Credit Tracing), a training-time interface that exposes executable or evidence backed atoms and traces them back to actions through dependency-valid proof edges. VICT redistributes group-relative advantage only along those edges, shifting credit assignment from rollout-side inference to verifierside tracing. It preserves the original terminal reward, abstains when evidence is incomplete or ambiguous, and changes only the training-time advantage tensor, requiring no learned critic, process labels, branch rollouts, or inference-time verifier access. On ALFWorld and WebShop, VICT improves substantially over outcome-only training and achieves strong performance alongside recent fine-grained credit methods; ablations rule out dense atom rewards, final-commit credit, temporal proximity, and sparsity as sufficient explanations.

Commentsaccepted by EMNLP2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑