布尔查询就足够了吗?
Boolean queries are all you need?
浏览论文内容
中文总结 AI 辅助
研究在TREC 2024 RAG赛道任务中,为基于大语言模型的搜索代理配备布尔检索引擎,仅依据语料库子串与查询匹配密度排名,无需监督学习等,结果表明简单模式匹配或足以用于智能搜索。
中文摘要 AI 辅助
我们为基于大语言模型的搜索代理配备了布尔检索引擎,以搜索TREC 2024 RAG赛道使用的MS MARCO V2.1去重片段集合。在86个主题的标准赛道子集中,在每个主题100次模型调用的预算下,该代理的NDCG@10为0.6863,高于许多密集、稀疏和学习稀疏的第一阶段检索器。排名仅基于与查询匹配的语料库子串的密度,无需监督学习、全局统计或词权重。形式上,查询语言表达了正则语言的一个严格子集,文档得分基于其包含的匹配项数量和长度。虽然结果更多是探索性而非确定性的,但表明简单模式匹配可能足以进行智能搜索。
英文摘要
We equipped an LLM-based search agent with access to a Boolean retrieval engine to search the MS MARCO V2.1 deduped segment collection used by the TREC 2024 RAG track. Over a standard track subset of 86 topics, and operating under a budget of 100 model calls/topic, the agent achieved an NDCG@10 of 0.6863, which would place it above many dense, sparse, and learned-sparse first-stage retrievers. Ranking is based solely on the density of corpus substrings matching a query, with no requirement for supervised learning, global statistics, or term weights. Formally, the query language expresses a strict subset of the regular languages, with a document's score based on the number and length of matches it contains. Although the results are more exploratory than definitive, because they are based on a single test collection that was publicly available during model training, they suggest that simple pattern matching may be sufficient for agentic search.