arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03177cs.AI

多样性并非歧义:面向开放域问答的准确高效歧义检测

Diversity is Not Ambiguity: Toward Accurate and Efficient Ambiguity Detection for Open-Domain QA

Jiwon Lee, Yong-chan Park, Jungin Hong, U Kang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对开放域问答中现有歧义检测方法的缺陷,提出ARCHIVE框架,构建QuireQA基准,在提升歧义检测准确率的同时大幅加快运行速度。

中文摘要 AI 辅助

问答(QA)系统如何判断一个查询是否存在歧义?歧义检测在开放域问答中至关重要,因为分类错误会导致回答错误的解读或进行不必要的澄清。然而,现有方法将答案多样性与歧义混为一谈,导致预测不准确;同时它们对所有查询采用统一处理,造成计算浪费。我们提出ARCHIVE(通过级联假设检验与冲突验证实现的歧义识别),这是一种准确高效的框架,通过逻辑冲突检测歧义:当一个查询的有效答案无法在单一解读下全部成立时,该查询即为歧义。ARCHIVE结合了用于检测表面可识别情况的轻量早退出编码器,以及对答案间逻辑关系进行建模的冲突推理模块,该模块由不变性目标强化,以提升对含噪答案集的鲁棒性。我们推出QuireQA,这是一个包含4703个查询的基准数据集,涵盖事实型、非事实型及格式错误的查询。实验表明,ARCHIVE的性能优于同类方法,其歧义F1值(F1-amb)提升最高达10.4%,非歧义F1值(F1-unamb)提升最高达21.6%,同时运行速度比最优同类方法快16倍。

英文摘要

How can question answering (QA) systems determine whether a query is ambiguous? Ambiguity detection is essential in open-domain QA, as misclassification leads to answering the wrong interpretation or unnecessary clarification. However, existing methods conflate answer diversity with ambiguity, leading to inaccurate predictions. They also process queries uniformly, resulting in wasteful computation. We propose ARCHIVE (Ambiguity Recognition via Cascaded Hypothesis Inspection and Conflict Verification), an accurate and efficient framework that detects ambiguity via logical conflict: a query is ambiguous when its valid answers cannot all be true under a single interpretation. ARCHIVE combines a lightweight early-exit encoder for surface-detectable cases with a conflict reasoning module that models logical relations among answers, reinforced by an invariance objective for robustness to noisy answer sets. We present QuireQA, a 4,703-query benchmark spanning factoid, non-factoid, and ill-formed queries. Experiments show ARCHIVE outperforms competitors, improving F1-amb by up to 10.4% and F1-unamb by up to 21.6%, while operating 16$\times$ faster than the best competitor.

补充信息

↑