HaPRL:面向视觉搜索智能体的人类锚定过程强化学习
HaPRL: Human-Anchored Process Reinforcement Learning for Visual Search Agent
- Tsinghua University(清华大学)
- Peng Cheng Laboratory(鹏城实验室)
- LMMs-Lab
- Beijing Institute of Technology(北京理工大学)
- University of New South Wales(新南威尔士大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
HaPRL提出首个以人类搜索行为锚定过程奖励的强化学习框架,通过细粒度人工标注和任务自适应评判,显著提升视觉搜索智能体训练效果,早期过程监督带来6.7倍改进。
AI中文摘要:
多轮视觉搜索智能体通过迭代决定查看位置来回答关于高分辨率图像的问题。针对这些智能体的强化学习仅对最终答案给予奖励,使得搜索过程不受监督。因此,推理过程错误但最终结果正确的错误路径频繁出现,这反过来导致训练效率低下,即沿着错误路径进行扩展。在本文中,我们引入了HaPRL,这是第一个利用人类搜索行为来强化搜索过程的框架。我们首先构建了一个标注平台,并收集了超过1000条带有细粒度行为信号的人工标注数据。在训练过程中,一个精心设计的评判器根据任务自适应权重对每次展开进行评分,并以人类标注者实际搜索同一图像的蒸馏轨迹为锚点。大量实验表明,HaPRL始终优于基于结果的强化学习,且早期阶段的过程监督在后续基于结果的扩展中带来了6.7倍的更大改进。我们的结果还证明了将模型行为与人类过程标注信号对齐的重要性,这为基础模型的训练提供了新的见解。
英文摘要:
Multi-turn visual search agents answer questions about high-resolution images by iteratively deciding where to look. Reinforcement learning for these agents rewards only the final answer, leaving the search process unsupervised. Consequently, faulty routes in which the reasoning process is erroneous yet the final result is correct arise frequently, which in turn leads to ineffective training, i.e., scaling along the wrong paths. In this paper, we introduce HaPRL, the first framework to reinforce the search process with human search behavior. We first build an annotation platform and collect 1K+ human-annotated data with fine-grained behavioral signals. During training, a carefully designed judge scores each rollout with task-adaptive weights, anchored on the distilled trace of how a human annotator actually searched the same image. Extensive experiments show that HaPRL consistently outperforms outcome-based RL, and early-stage process supervision yields 6.7x more improvement in subsequent outcome-based scaling. Our results also demonstrate the importance of aligning model behavior with human process annotation signals, which offer new insight into the training of foundation models.