ASCT:用于智能体强化学习中信用分配的对抗树注意力搜索
ASCT: Attentive Search over Counterfactual Trees for Credit Assignment in Agentic Reinforcement Learning
- University of Hong Kong(香港大学)
- The Chinese University of Hong Kong(香港中文大学)
- Jiangxi Science and Technology Normal University(江西科技师范大学)
- Shenzhen University(深圳大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
ASCT通过对抗树搜索将训练时的多步搜索转化为局部动作信用,提升智能体强化学习的信用分配,在HotpotQA上优于现有方法,并减少计算开销。
中文摘要 AI 辅助
终端效用评估完整的智能体工作流,但学习需要对工作流内部决策的信用分配。我们引入了对抗树注意力搜索(ASCT),这是一个将训练时的多步搜索转化为局部动作信用的框架。在演员访问的状态下,辅助树评估来自相同可恢复前缀的替代合法动作。其动作价值表以冻结的演员概率为中心,并为演员采样的轨迹上的PPO提供信用。该协议将反事实评估与策略学习联系起来,同时仅部署演员。均匀、UCT和成本感知的AgentUCT实例化了该框架。在HotpotQA智能体检索增强生成上,三种方法均优于轨迹回报PPO和工作流适配的VinePPO的平均保留效用。在三个种子中,ASCT-AgentUCT达到0.6187效用,而VinePPO为0.5939,在答案F1和执行成本上有所提升,并减少了50.3%的记录辅助Qwen令牌。迁移和组件描述研究考察了训练设置之外的习得策略。
英文摘要
Terminal utility evaluates a complete agentic workflow, but learning requires credit for the decisions within it. We introduce Attentive Search over Counterfactual Trees (ASCT), a framework that turns training-time multi-step search into local action credit. At actor-visited states, an auxiliary tree evaluates alternative legal actions from the same recoverable prefix. Its action-value table is centered by the frozen actor's probabilities and supplies credit for PPO on actor-sampled trajectories. This protocol connects counterfactual evaluation to policy learning while deploying the actor alone. Uniform, UCT, and cost-aware AgentUCT instantiate the framework. On HotpotQA agentic retrieval-augmented generation, all three improve mean held-out utility over trajectory-return PPO and workflow-adapted VinePPO. Across three seeds, ASCT-AgentUCT reaches 0.6187 utility versus 0.5939 for VinePPO, with gains in answer F1 and execution cost, and uses 50.3% fewer recorded auxiliary Qwen tokens. Transfer and component-description studies examine the learned policies beyond the training setting.