arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35215cs.AI

ASCT:用于智能体强化学习中信用分配的对抗树注意力搜索

ASCT: Attentive Search over Counterfactual Trees for Credit Assignment in Agentic Reinforcement Learning

  • University of Hong Kong(香港大学)
  • The Chinese University of Hong Kong(香港中文大学)
  • Jiangxi Science and Technology Normal University(江西科技师范大学)
  • Shenzhen University(深圳大学)

机构由 AI 辅助整理,请以论文原文为准。

Yang Li, Jinhan Yang, hai liu, Di Wan, Xiyu Chen, Zongsi Xu, Tuo Zhou, Sheng Zhong, Sergey Volkov, Ye Luo, Hao Sun

中文总结 AI 辅助

ASCT通过对抗树搜索将训练时的多步搜索转化为局部动作信用,提升智能体强化学习的信用分配,在HotpotQA上优于现有方法,并减少计算开销。

中文摘要 AI 辅助

终端效用评估完整的智能体工作流,但学习需要对工作流内部决策的信用分配。我们引入了对抗树注意力搜索(ASCT),这是一个将训练时的多步搜索转化为局部动作信用的框架。在演员访问的状态下,辅助树评估来自相同可恢复前缀的替代合法动作。其动作价值表以冻结的演员概率为中心,并为演员采样的轨迹上的PPO提供信用。该协议将反事实评估与策略学习联系起来,同时仅部署演员。均匀、UCT和成本感知的AgentUCT实例化了该框架。在HotpotQA智能体检索增强生成上,三种方法均优于轨迹回报PPO和工作流适配的VinePPO的平均保留效用。在三个种子中,ASCT-AgentUCT达到0.6187效用,而VinePPO为0.5939,在答案F1和执行成本上有所提升,并减少了50.3%的记录辅助Qwen令牌。迁移和组件描述研究考察了训练设置之外的习得策略。

英文摘要

Terminal utility evaluates a complete agentic workflow, but learning requires credit for the decisions within it. We introduce Attentive Search over Counterfactual Trees (ASCT), a framework that turns training-time multi-step search into local action credit. At actor-visited states, an auxiliary tree evaluates alternative legal actions from the same recoverable prefix. Its action-value table is centered by the frozen actor's probabilities and supplies credit for PPO on actor-sampled trajectories. This protocol connects counterfactual evaluation to policy learning while deploying the actor alone. Uniform, UCT, and cost-aware AgentUCT instantiate the framework. On HotpotQA agentic retrieval-augmented generation, all three improve mean held-out utility over trajectory-return PPO and workflow-adapted VinePPO. Across three seeds, ASCT-AgentUCT reaches 0.6187 utility versus 0.5939 for VinePPO, with gains in answer F1 and execution cost, and uses 50.3% fewer recorded auxiliary Qwen tokens. Transfer and component-description studies examine the learned policies beyond the training setting.

补充信息

↑