arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34805cs.AI

SIPO:面向树结构智能体强化学习的选择性推断策略优化

SIPO: Selective-Inference Policy Optimization for Tree-Structured Agentic RL

Zenghuang Fu, Ningqi Chen, Mingda Jia, Xiaofeng Han, Zhaoyang Li, Qiuyuan Ai, Zelong Zheng, Haoyu Wu, Tianyu Fu, Chenxu Zhao, Minghui Wu, Guannan He, Changwei Wang

首次发表
浏览论文内容

中文总结 AI 辅助

SIPO提出选择性推断策略优化,通过无标度分支准则、可交换分支和顺序统计量校正,解决树结构强化学习中自适应扩展导致的统计不对称性,在七个QA基准上取得领先性能。

中文摘要 AI 辅助

树结构强化学习通过比较不同的后续路径并将终端奖励传播到中间决策来训练搜索智能体。然而,自适应扩展造成了一种统计上的不对称性:现有分支是根据其自身的生成统计量被选中的,而新的兄弟分支则是在选择之后才被采样。当该统计量与回报相关联时,分支值可能既反映选择历史,也反映延续质量,即使它们共享同一个父节点。我们提出了选择性推断策略优化(SIPO),将这一区别纳入基于树的信用分配中。其无标度分支准则使生成分数和兄弟惩罚保持在一致的相对尺度上;可交换分支为每个选中的父节点提供多个新的延续;顺序统计量校正则利用选择排名和估计的分数-结果关联来调整保留的现有分支值。这些机制保持了叶子预算和宿主策略优化目标不变。在使用Qwen3-4B、Qwen3-8B和Qwen2.5-7B的七个问答基准上,SIPO在比较方法中取得了最高的多跳和单跳平均分数。在Qwen3-8B上,它分别比AT²PO提高了1.31和1.07个百分点,并在七个基准中的六个上排名第一。组件消融实验评估了各个变更及其组合效果,而早期训练配对诊断显示,选中-新分支值之间存在差距,而新-新参考值接近零。这些结果共同支持在构建和评估搜索智能体轨迹时考虑选择历史。我们的代码可在该https URL获取。

英文摘要

Tree-structured reinforcement learning trains search agents by comparing alternative continuations and propagating terminal rewards to intermediate decisions. Adaptive expansion, however, creates a statistical asymmetry: an incumbent is selected using its own generation statistic, whereas fresh siblings are sampled after selection. When that statistic is associated with return, branch values can reflect selection history as well as continuation quality, even for a shared parent. We propose Selective-Inference Policy Optimization (\SIPO{}), which incorporates this distinction into tree-based credit estimation. Its scale-free branch criterion keeps generation scores and sibling penalties on a consistent relative scale; exchangeable branching supplies multiple fresh continuations from each selected parent; and order-statistic correction adjusts retained incumbent values using selection rank and the estimated score--outcome association. These mechanisms preserve the leaf budget and the host policy optimisation objective. Across seven QA benchmarks using Qwen3-4B, Qwen3-8B, and Qwen2.5-7B, \SIPO{} achieves the highest reported multi-hop and single-hop averages among the compared methods. On Qwen3-8B, it improves these averages over AT\textsuperscript{2}PO by $1.31$ and $1.07$ percentage points, respectively, and ranks first on six of seven benchmarks. Component ablations evaluate the individual and combined changes, while early-training paired diagnostics show a selected--fresh value gap alongside a near-zero fresh--fresh reference. Together, these results support accounting for selection history when constructing and evaluating search-agent rollouts. Our code is available at https://github.com/Zenghuang-Fu/SIPO

发表机构

  • University of Chinese Academy of Sciences(中国科学院大学)
  • Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
  • The University of Hong Kong(香港大学)
  • Peking University(北京大学)
  • Mininglamp Technology(明略科技)
  • Key Laboratory of Computing Power Network and Information Security, Ministry of Education(计算力网络与信息安全教育部重点实验室;齐鲁工业大学(山东省科学院)山东省计算中心)
  • Shandong Computer Science Center, Qilu University of Technology (Shandong Academy of Sciences)(计算力互联网与服务计算重点实验室,山东省计算机科学基础研究中心)
  • Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science

机构由 AI 辅助整理,请以论文原文为准。

↑