BiCAA:面向搜索增强智能体的双向信用分配
BiCAA: Bidirectional Credit Assignment for Search-Augmented Agent
AI总结:
针对搜索增强智能体多步搜索任务中普通GRPO的稀疏监督问题,提出BiCAA双向信用分配框架,融合前向可解性增益与后见成功关键性构建密集过程奖励,实现稳定策略优化并取得有竞争力的QA性能。
AI中文摘要:
多步搜索是搜索智能体的核心能力,使其能迭代获取、优化并整合外部证据以解决复杂推理问答(QA)问题。然而,普通GRPO仅基于模型最终输出分配奖励,仅提供结果监督,无中间推理步骤的监督信号,这种稀疏监督易导致训练不稳定及多步搜索任务中的冗余搜索行为。为缓解此局限,我们采用过程奖励提供分步监督信号。针对该过程奖励,我们提出两个互补标准判断每个搜索步骤:该步骤是否产生新证据以推进问题解决,是否在整体推理轨迹中形成高效、关键的中间决策。基于此,我们提出BiCAA:一种为搜索增强智能体提供密集、差异化过程奖励的双向信用分配框架。BiCAA通过融合两个互补信号构建双向过程奖励:前向可解性增益与后见成功关键性。前者量化答案合理性的分步改进,后者通过基于后见结果的关键性评分评估每个步骤对最终成功的必要性。我们对两个信号进行调制与聚合,再与结果奖励融合。在搜索增强QA基准上的实验表明,BiCAA可稳定策略优化、减少冗余搜索行为并取得有竞争力的性能。
英文摘要:
Multi-step search is a fundamental capability for search agents, enabling them to iteratively acquire, refine, and integrate external evidence for complex reasoning QA. However, vanilla GRPO allocates rewards exclusively based on the model's final outputs, yielding outcome-only supervision with no supervisory signals for intermediate reasoning steps. Such sparse supervision easily causes training instability and redundant search behaviors on multi-step search tasks. To mitigate this limitation, we adopt process reward to deliver stepwise supervision signals. For this process reward, we propose two complementary criteria to judge each search step: whether the step yields new evidence to facilitate problem solving, and whether it forms an efficient, pivotal intermediate decision within the overall reasoning trajectory. Building on this insight, we propose BiCAA: a bidirectional credit assignment framework that delivers dense, distinguishing process rewards for search-augmented agents. BiCAA builds bidirectional process rewards by fusing two complementary signals: forward solvability gain and hindsight success criticality. The former quantifies step-wise improvements in answer plausibility, while the latter evaluates each step's necessity for final success via hindsight outcome-based criticality scoring. We modulate and aggregate the two signals and then fuse them with the outcome reward. Experiments on search-augmented QA benchmarks show that BiCAA stabilizes policy optimization, reduces redundant search behavior, and achieves competitive performance.