arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2409.18575cs.IR

语料库知情的澄清问题检索增强生成

Corpus-informed Retrieval Augmented Generation of Clarifying Questions

Antonios Minas Krasakis, Andrew Yates, Evangelos Kanoulas

更新

AI总结:

本研究开发基于RAG的语料库知情澄清问题生成模型,发现现有数据集搜索意图与语料库不匹配导致幻觉,提出数据集增强方法对齐真实澄清与语料库,并呼吁构建支持语料库知情澄清的数据。

AI中文摘要:

本研究旨在开发为网络搜索生成语料库知情澄清问题的模型,以确保问题与检索语料库中的可用信息对齐。我们展示了检索增强语言模型(RAG)在此过程中的有效性,强调其能够(i)联合建模用户查询和检索语料库以精准定位不确定性并端到端地请求澄清,以及(ii)建模更多证据文档,这可用于增加所提问题的广度。然而,我们观察到在当前数据集中,搜索意图很大程度上不受语料库支持,这对训练和评估均有问题。这导致问题生成模型产生“幻觉”,即建议语料库中不存在的意图,可能对性能产生不利影响。为解决此问题,我们提出数据集增强方法,将真实澄清与检索语料库对齐。此外,我们探索在推理过程中提升证据池相关性的技术,但发现识别语料库内的真实意图仍具挑战。我们的分析表明,该挑战部分源于当前数据集对澄清分类法的偏向,并呼吁需要能支持生成语料库知情澄清的数据。

英文摘要:

This study aims to develop models that generate corpus informed clarifying questions for web search, in a way that ensures the questions align with the available information in the retrieval corpus. We demonstrate the effectiveness of Retrieval Augmented Language Models (RAG) in this process, emphasising their ability to (i) jointly model the user query and retrieval corpus to pinpoint the uncertainty and ask for clarifications end-to-end and (ii) model more evidence documents, which can be used towards increasing the breadth of the questions asked. However, we observe that in current datasets search intents are largely unsupported by the corpus, which is problematic both for training and evaluation. This causes question generation models to ``hallucinate'', ie. suggest intents that are not in the corpus, which can have detrimental effects in performance. To address this, we propose dataset augmentation methods that align the ground truth clarifications with the retrieval corpus. Additionally, we explore techniques to enhance the relevance of the evidence pool during inference, but find that identifying ground truth intents within the corpus remains challenging. Our analysis suggests that this challenge is partly due to the bias of current datasets towards clarification taxonomies and calls for data that can support generating corpus-informed clarifications.

↑