为什么要枚举而非采样?基因组工具选择的精确策略优化
Why Sample What You Can Enumerate? Exact Policy Optimization for Genomic Tool Selection
- School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)理工学院)
- Shenzhen Future Network of Intelligence Institute (FNii-Shenzhen)(深圳市未来智联网络研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对基因组工具选择,提出FGPO精确枚举所有工具子集并优化期望,替代GRPO采样近似,在15个设置中平均提升6.75分并减少推理调用。
AI中文摘要:
在冻结推理器上进行强化学习已成为教导策略调用哪些外部工具的常见方法。我们表明,在专业科学环境中,当完整的工具子集空间可枚举时,这种方法在结构上变得不匹配。在那里,一小部分反复出现的计算能力覆盖了整个领域,因此工具子集的空间是组合性的,但小到足以枚举,而GRPO仍然从少量采样的轨迹中估计动作期望。更糟糕的是,随着训练的成功,这种近似会退化:当策略集中于偏好的子集时,它会重新采样这些子集,采样的奖励发生冲突,组归一化的优势消失。在基因组推理中,没有奖励信号的问题比例从统一参考策略下的0.2%上升到GRPO训练后的20.8%。作为补救措施,我们引入了FGPO(全组策略优化),它(1)对每个工具子集进行评分并优化精确的动作期望,因此每次更新都能看到完整的动作空间,并且(2)将每个问题-子集对的奖励预计算到一个穷举表中,完全从训练循环中移除冻结推理器的调用。在五个冻结推理器和三个基因组基准上,FGPO在所有15种设置中平均优于GRPO 6.75分,最高达14.20分,而标准的按需GRPO调度需要2.4倍的冻结推理器奖励评估,并且在GenomeQA上,FGPO将每个问题调用的工具数从2.36减少到1.40。
英文摘要:
Reinforcement learning over a frozen reasoner has become a common recipe for teaching a policy which external tools to invoke. We show that this recipe becomes structurally mismatched in specialist scientific settings where the complete tool-subset space is enumerable. There, a small set of recurring computational capabilities covers the domain, so the space of tool subsets is combinatorial yet small enough to enumerate, and GRPO still estimates an action expectation from a handful of sampled rollouts. Worse, the approximation degrades as training succeeds: as the policy concentrates on preferred subsets it resamples them, sampled rewards collide, and the group-normalized advantage vanishes. On genomic reasoning the fraction of questions yielding no reward signal rises from 0.2% under a uniform reference policy to 20.8% after GRPO training. As a remedy, we introduce FGPO (Full-Group Policy Optimization), which (1) scores every tool subset and optimizes the exact action expectation, so each update sees the complete action space, and (2) precomputes the reward of each question--subset pair into an exhaustive table, removing frozen-reasoner calls from the training loop entirely. Across five frozen reasoners and three genomic benchmarks, FGPO outperforms GRPO in all 15 settings by 6.75 points on average and up to 14.20, while a standard on-demand GRPO schedule would require 2.4 times as many frozen-reasoner reward evaluations and, on GenomeQA, FGPO cuts invoked tools per question from 2.36 to 1.40.