arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于语料反馈的多步查询重写用于对话式搜索

Learning Multi-Step Query Rewriting via Corpus Feedback for Conversational Search

João Coelho, Hong Wang, Jie Yuan, Zhuoer Wang, Samson Koelle, Wei Niu

arXiv 2610.10955首次发表:更新:

发表机构

INESC-ID; Carnegie Mellon University; Amazon(INESC-ID; 卡内基梅隆大学; 亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究将对话式查询重写转化为序列检索问题,提出基于语料反馈的多步重写智能体,无需人工标注,在多个数据集上优于基线且可泛化,提升了检索效果。

AI 中文摘要

对话式查询重写(Conversational Query Rewriting,CQR)将依赖上下文的用户轮次转换为检索器可用的独立查询,现有多数方法仅从对话历史中执行单步重写,在获取任何语料证据以修正指代消解或词汇问题前就已确定重写结果。本文将CQR重新定义为序列检索问题:智能体重写当前轮次、执行检索,并基于返回的段落调整下一次重写。智能体在包含三类重写操作的类型空间中行动,包括将对话意图解析为独立查询、生成词汇改写、合成伪文档用于文档间匹配,以及结束回合的停止动作。我们采用监督微调后再用强化学习训练策略,仅使用单一检索质量奖励,无需人工重写标注。在TopiOCQA和QReCC数据集上,该智能体优于多个检索对齐基线,且在不同检索后端均保持有效,无需额外训练即可泛化到CAsT基准。进一步分析显示,仅通过检索奖励优化,学习到的策略会将伪文档锚定到前期步骤检索到的段落,大幅提升检索效果。

英文摘要

Conversational Query Rewriting (CQR) turns a context dependent user turn into a standalone query for a retriever, and most methods do this in a single step from the dialogue history before retrieving once. The rewrite is therefore fixed before any corpus evidence is available to correct its reference resolution or its vocabulary. We recast CQR as a sequential retrieval problem: an agent rewrites the current turn, retrieves, and conditions its next rewrite on the returned passages. The agent acts in a typed space of three rewriting operations, resolving conversational intent into a standalone query, generating lexical reformulations, or synthesizing pseudo-documents for document-to-document matching, together with a stop action that ends the episode. We train the policy with supervised fine-tuning followed by reinforcement learning against a single retrieval-quality reward, using no human rewrite annotations. Across TopiOCQA and QReCC, the agent outperforms several retrieval-aligned baselines, while remaining effective across retrieval backends and generalizing to the CAsT benchmarks without additional training. Further analysis shows that, through retrieval-reward optimization alone, the learned policy develops a behavior of grounding pseudo-documents in passages retrieved by earlier steps, substantially improving retrieval.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑