arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

先澄清再搜索:面向端到端 nugget 恢复的深度搜索澄清基准

Clarify-Then-Search: A Clarification Benchmark for Deep Search with End-to-End Nugget Restoration

Deqiang Huang, Jingbo Zhou, Xinjiang Lu, Tong Xu, Hua Wu, Enhong Chen

arXiv 2608.20357首次发表:更新:

发表机构

University of Science and Technology of China; Baidu Inc.(中国科学技术大学; 百度公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出 Clarify-Then-Search 基准,评估 LLM 生成的澄清问题对深度搜索效用的提升,实验显示 k=1 时澄清优于基线,GPT-5.2 和 ERNIE-4.5-Turbo-128K 分别在 k=1、k=3 时表现最佳。

AI 中文摘要

深度搜索对用户表述不明确的查询较为脆弱:缺失时间、位置、范围或定义等约束会导致检索漂移和答案不完整。我们提出 Clarify-Then-Search,这是一个用于评估大语言模型(LLM)生成的澄清问题是否能提升下游深度搜索效用的基准。该基准基于百度搜索引擎的真实查询数据构建,包含 518 个精心整理的实例,每个实例包含一个意图查询和对应的表述不明确的查询。对于每个意图查询,我们运行 WebDancer 一次以归档证据,并构建静态黄金参考,作为带有可追溯源标识符的、基于证据的加权 nugget。在评估阶段,澄清器(Clarifier)会提出 k 个问题,k 取值为 {1,2,3};闭卷用户回答器(User Answerer)仅回复意图查询中明确陈述的信息,否则返回未知;闭卷重写器(Rewriter)仅使用表述不明确的查询和已获取的问答对生成重写后的查询。随后 WebDancer 对重写后的查询执行搜索,我们使用 restore_score_100 对端到端效用进行评分,这是一种针对静态黄金的、带有部分信用的加权 nugget 召回率评分。在所有评估的模型中,k=1 时澄清比无交互基线有所提升,更大的预算通常会带来进一步的提升。GPT-5.2 在 k=1 时取得最高平均分数,而 ERNIE-4.5-Turbo-128K 在 k=3 时成为整体表现最佳的模型。诊断结果显示存在一致的失败模式:许多系统过度提出仅关于区域的问题,而这类问题通常无法从意图查询中得到解答,从而返回未知。Clarify-Then-Search 支持对先澄清再搜索的管道进行抗泄露且可复现的评估,可对深度搜索中问题效用、可回答性和预算效应进行细粒度分析。

英文摘要

Deep search is brittle on underspecified user queries: missing constraints such as time, location, scope, or definitions can lead to retrieval drift and incomplete answers. We introduce Clarify-Then-Search, a benchmark for evaluating whether LLM-generated clarification questions improve downstream deep-search utility. Built on real-world query data from the Baidu search engine, the benchmark contains 518 curated instances, each with an intent query and a corresponding underspecified query. For each intent query, we run WebDancer once to archive evidence and construct a static golden reference as weighted, evidence-grounded nuggets with traceable source identifiers. At evaluation time, a Clarifier asks k in {1, 2, 3} questions; a closed-book User Answerer replies only with information explicitly stated in the intent query, otherwise returning unknown; and a closed-book Rewriter produces a rewritten query using only the underspecified query and the elicited question-answer pairs. WebDancer then executes on the rewritten query, and we score end-to-end utility using restore_score_100, a weighted nugget-recall score with partial credit against the static gold. Across all evaluated models, clarification improves over the no-interaction baseline at k=1, and larger budgets generally yield further gains. GPT-5.2 achieves the highest mean score at k=1, while ERNIE-4.5-Turbo-128K becomes the overall top-performing model at k=3. Diagnostics reveal a consistent failure mode: many systems over-ask region-only questions that are often unanswerable from the intent and thus elicit unknown. Clarify-Then-Search enables leakage-resistant and reproducible evaluation of clarify-then-search pipelines, with fine-grained analyses of question utility, answerability, and budget effects in deep search.

CommentsAccepted to KDD 2026 Datasets and Benchmarks Track. 12 pages, 4 figures, 11 tables

DOI:10.1145/3770855.3817586

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑