arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14646cs.CLcs.AI

CompCQR:用于免训练对话式搜索的组合式查询生成

CompCQR: Compositional Query Generation for Training-Free Conversational Search

  • Seoul National University(首尔大学)
  • LG AI Research(LG AI研究院)

机构由 AI 辅助整理,请以论文原文为准。

Yunah Jang, Kang-il Lee, Joongbo Shin, Kyomin Jung

AI总结:

针对对话式搜索中查询歧义与检索器不对齐问题,提出免训练的组合式查询生成方法CompCQR,以最少LLM调用生成大量查询并构建高质量文档集,在四个基准上MRR相对提升高达22.5%。

AI中文摘要:

在信息寻求场景中,与大型语言模型(LLM)的多轮交互正变得越来越普遍。然而,用户查询往往具有歧义且依赖上下文,这使得它们不适合直接用作检索器查询。对话式查询重构(CQR)通过将当前话语重写为基于对话历史的独立查询来解决这一问题。近期基于LLM的CQR方法取得了强劲性能;然而,它们反复调用LLM以及与下游检索器不对齐的问题仍然是挑战。在这项工作中,我们从观察出发,即检索器对内容排序高度敏感:简单地重新排列相同内容就可能导致检索覆盖范围和性能的变化。基于此,我们提出了一种新颖的免训练方法,通过组合性地结合一小组原子组件,以最少的LLM使用量生成大量查询。我们进一步应用LLM推理来构建一个高质量的文档集,该文档集在捕捉用户核心意图的同时平衡精确率和召回率。我们的框架可泛化到开源和闭源LLM以及稠密和稀疏检索器。它在四个广泛使用的对话基准上取得了强劲性能,与之前最先进的基线相比,相对MRR提升了高达22.5%,同时LLM调用次数大大减少。

英文摘要:

Multi-turn interactions with LLMs are becoming increasingly common in information-seeking scenarios. However, user queries are often ambiguous and context-dependent, making them ill-suited for direct use as retriever queries. Conversational query reformulation (CQR) addresses this issue by rewriting the current utterance into a stand-alone query grounded in the dialogue history. Recent LLM-based CQR approaches achieve strong performance; however, their repeated LLM invocations and misalignment with downstream retrievers remain challenges. In this work, we begin from the observation that retrievers are highly sensitive to content ordering: simply reordering the same content can lead to changes in retrieval coverage and performance. Based on this, we propose a novel training-free method that generates a very large number of queries with minimal LLM usage by compositionally combining a small set of atomic components. We further apply LLM reasoning to construct a high-quality document set that balances precision and recall while capturing the user's core intent. Our framework generalizes across both open- and closed-source LLMs as well as dense and sparse retrievers. It achieves strong performance on four widely used conversational benchmarks, with up to 22.5% relative MRR improvement over the previous state-of-the-art baseline with far fewer LLM calls.

补充信息

↑