arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37749cs.IR

迈向半自动比较基于关键词与语义搜索的准确性

Towards Semi-Automatically Comparing Keyword-Based and Semantic Search Accuracy

  • Technical University of Munich(慕尼黑工业大学)
  • Gofore GmbH(Gofore有限公司)
  • University of Stuttgart(斯图加特大学)

机构由 AI 辅助整理,请以论文原文为准。

Mohamed Ben Salha, Fiete Lüer, Maik Betka, Stefan Wagner

AI总结:

针对关键词与语义搜索缺乏定量比较的问题,提出半自动框架,通过等价类评估排序准确性与信息完整性,并经工业案例验证语义搜索优势。

AI中文摘要:

信息检索(IR)在管理大型数据集中的重要性日益凸显,这揭示了传统基于关键词的搜索系统的显著局限性。近年来,出现了诸如检索增强生成(RAG)等上下文感知的基于聊天的搜索方法,但将其与基于关键词的系统进行比较的评估往往依赖于主观的用户反馈。这两种范式之间缺乏严谨的定量比较。本研究引入了一个新颖的初步框架,用于定量评估产生不同输出格式(如列表和消息)的搜索系统的IR准确性。该框架聚焦于两个关键方面:基于关键词的系统的排序准确性,以及基于语义的聊天系统的检索信息完整性。我们的方法通过针对特定领域上下文(如公司或问题)定制的可互换等价类,实现了对语义和基于关键词方法的半自动比较。我们通过一项工业案例研究验证了该框架,证明了上下文感知搜索相对于基于关键词的方法具有统计学上显著的改进,并得到了包括Mann-Whitney U检验在内的分析支持。凭借其适应性设计,所提出的框架为客观评估基于关键词和基于语义的聊天搜索方法提供了坚实的基础。

英文摘要:

The increasing importance of Information Retrieval (IR) in managing large datasets has highlighted significant limitations in traditional keyword-based search systems. Context-aware chat-based search methods, such as Retrieval Augmented Generation (RAG), have recently emerged, but their evaluation compared to keyword-based systems often relies on subjective user feedback. A rigorous, quantitative comparison between these paradigms remains lacking. This work introduces a novel, preliminary framework to quantitatively assess IR accuracy of search systems that produce different output formats, such as lists and messages. It focuses on two key aspects: the ranking accuracy for keyword-based systems and the completeness of retrieved information for semantic chat-based systems. Our approach enables semi-automatic comparisons of semantic and keyword-based methods using interchangeable equivalence classes tailored to domain-specific contexts (e.g., companies or problems). We validate the framework through an industrial case study, demonstrating statistically significant improvements in context-aware search over keyword-based methods, supported by analyses including the Mann-Whitney U-Test. With its adaptable design, the proposed framework provides a strong foundation for objectively assessing keyword-based and semantic chat-based search methods.

补充信息

↑