arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BOUND:在搜索-控制边界处的摘要引导式校正偏好蒸馏

BOUND: Brief-Guided Corrective Preference Distillation at Search-Control Boundaries

Qingying Niu, Ruiyang Ren, Wayne Xin Zhao, Yaliang Li

arXiv 2608.08768首次发表:更新:

AI 中文总结

针对大语言模型深度搜索智能体的持续漂移问题,提出BOUND框架,通过摘要引导构建偏好对,用DPO蒸馏偏好,在7个基准上多数数据集和指标优于基线。

AI 中文摘要

基于大语言模型(LLM)的深度搜索智能体通过迭代检索与推理解决任务,但局部相关证据会引发持续的错误锚点漂移、约束漂移或局部主题漂移。现有方法对轨迹、结果或步骤进行监督,但很少区分与任务对齐的延续和会加剧漂移的局部看似合理的延续。我们提出BOUND,一种用于解决持续搜索漂移的摘要引导式校正偏好蒸馏框架。对于每个学生智能体在决策时的状态,BOUND构建教师侧的搜索状态摘要,该摘要在总结已确认证据、缺失信息和漂移状态的同时,保留原始搜索目标和关键约束。在该摘要的引导下,教师判断学生的延续是否包含可能影响后续决策的可校正局部搜索-控制错误。结合rollout(滚动)结果,此判断决定是构建学生特定校正与原始延续之间的校正对比,还是构建支持答案与不必要检索延续之间的终止对比。每个经验证的状态匹配偏好对都对应一个搜索-控制边界。直接偏好优化(DPO)将这些偏好蒸馏到学生智能体中,而摘要和教师侧计算仅在训练阶段使用。我们在四个多跳QA基准和三个深度搜索基准上评估BOUND。在我们重新运行基线的六个基准中,BOUND在五个数据集和14个指标中的12个上表现领先。在相同的搜索-控制接口和匹配设置下,BOUND在Bamboogle上比Trajectory SFT高出5.6个精确匹配(EM)点,在BrowseComp-Plus上高出4.8个准确率点。代码可在此https URL获取。

英文摘要

Large language model (LLM)-based deep search agents solve tasks through iterative retrieval and reasoning, but locally relevant evidence can cause persistent wrong-anchor drift, constraint drift, or local-topic drift. Existing methods supervise trajectories, outcomes, or steps, but rarely distinguish task-aligned continuations from locally plausible ones that reinforce drift. We propose BOUND, a brief-guided corrective preference distillation framework for persistent search drift. For each student-induced decision-time state, BOUND constructs a teacher-side search-state brief that preserves the original search target and key constraints while summarizing confirmed evidence, missing information, and drift status. Guided by the brief, the teacher determines whether the student's continuation contains a correctable local search-control error likely to affect subsequent decisions. Together with the rollout outcome, this assessment determines whether to construct a corrective contrast between a student-specific correction and the original continuation, or a termination contrast between a supported answer and an unnecessary retrieval continuation. Each validated state-matched preference pair operationalizes a search-control boundary. Direct preference optimization (DPO) distills these preferences into the student, while the brief and teacher-side computation remain confined to training. We evaluate BOUND on four multi-hop QA benchmarks and three deep-search benchmarks. Across the six benchmarks for which we reran baselines, BOUND leads on five datasets and 12 of 14 metrics. Under the same search-control interface and matched settings, BOUND outperforms Trajectory SFT by 5.6 EM points on Bamboogle and 4.8 accuracy points on BrowseComp-Plus. Code is available at https://github.com/RUCAIBox/BOUND.

Comments15 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑