TRCA:面向长 horizon 大语言模型智能体的逐步规则信用分配
TRCA: Transition-wise Rubric Credit Assignment for Long-horizon LLM Agents
AI总结:
针对长 horizon LLM 智能体稀疏最终结果优化导致的细粒度信用分配难题,提出无需学习评估器或成功锚点的 TRCA,在多基准上取得显著性能提升。
AI中文摘要:
长 horizon 大语言模型(LLM)智能体通常通过稀疏的最终结果进行优化,这使得多步交互中的细粒度信用分配变得困难。现有方法要么依赖过程评估器,会产生标注和推理成本,要么从成功轨迹中推导步级信用。然而,在强化学习初期,成功轨迹极为稀缺,这大幅削弱了基于锚点的方法。我们提出逐步规则信用分配(TRCA),该方法无需学习评估器或成功锚点,直接从动作诱导的转换中推导步级监督信号。TRCA 使用证据、执行和无效性规则评估每个转换,以捕捉与任务相关的信息获取、有效任务执行以及无效或倒退行为。基于这些判断,基础规则奖励衡量局部转换质量,突破规则奖励追踪新覆盖的证据和执行条件,以奖励增量任务进展。结合最终结果,这些信号为策略优化提供细粒度的步级优势。在 ALFWorld、WebShop 和七个搜索增强问答基准上的实验显示,TRCA 相比评估的基线方法取得了一致提升。使用 Qwen2.5-7B-Instruct 时,TRCA 将 WebShop 得分提升了 6.0%-12.6%;使用 Qwen2.5-3B-Instruct 时,将平均 SearchQA 得分提升了 1.9%-18.3%。这些结果证明了逐步规则信用分配在稀疏成功锚点的长 horizon 任务中的有效性。
英文摘要:
Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, making fine-grained credit assignment across multi-step interactions difficult. Existing approaches either rely on process evaluators, which incur annotation and inference costs, or derive step-level credit from successful trajectories. However, successful trajectories are extremely scarce during early-stage reinforcement learning, substantially weakening anchor-based methods. We propose Transition-wise Rubric Credit Assignment (TRCA), which derives step-level supervision directly from action-induced transitions without learned evaluators or successful anchors. TRCA evaluates each transition using Evidence, Execution, and Invalidity rubrics to capture task-relevant information acquisition, valid task execution, and invalid or regressive behavior. From these judgments, Foundational Rubric Reward measures local transition quality, while Breakthrough Rubric Reward tracks newly covered Evidence and Execution conditions to reward incremental task progress. Combined with terminal outcomes, these signals produce fine-grained step-level advantages for policy optimization. Experiments on ALFWorld, WebShop, and seven search-augmented question-answering benchmarks show consistent improvements over the evaluated baselines. With Qwen2.5-7B-Instruct, TRCA improves the WebShop score by 6.0%-12.6%; with Qwen2.5-3B-Instruct, it improves the average SearchQA score by 1.9%-18.3%. These results demonstrate the effectiveness of transition-wise rubric credit assignment for long-horizon tasks with sparse successful anchors.