发表机构
University of Glasgow(格拉斯哥大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探究IR模型对集合增长的鲁棒性,将模型分为MDA和MDD两类,发现两类模型均无法完全抵御非相关文档添加带来的性能下降,且MDA检索器效果优于MDD,二者重排序器效果相当。
AI 中文摘要
信息检索(IR)系统旨在在一个文档集合中识别相关文档。在实际应用中,文档集合是动态的,会频繁添加新文档。我们认为,理想情况下,当向集合中添加非相关文档时,检索器的有效性不应下降。本研究将这一概念形式化,并通过合并两个主题重叠可忽略的集合进行实证评估。我们假设,IR模型对集合中其他文档进行排序时的条件设定方式(例如BM25中的IDF组件或列表式重排序器中的上下文文档),对其应对非相关文档添加的鲁棒性起着重要作用。我们将模型大致分为不依赖其他文档的模型(多文档无关,MDA)和依赖其他文档的模型(多文档相关,MDD)。我们的结果显示,无论是MDD还是MDA模型,在添加非相关文档时都并非完全鲁棒,所有模型均表现出一定的性能下降。有趣的是,在我们测试的模型中,MDA检索器的效果优于MDD,而MDD和MDA重排序器的效果相当。
英文摘要
Information Retrieval (IR) systems seek to identify relevant documents within a collection. In practical applications, collections are dynamic, with documents frequently added. We argue that ideally, a retriever's effectiveness should not decrease when non-relevant documents are added to a collection. This study formalises this concept and empirically evaluates it by merging two collections with negligible topic overlap. We hypothesise that the way an IR model conditions its ranking on other documents in a collection (e.g., the IDF component in BM25 or contextual documents in listwise rerankers) plays an important role in its robustness to the addition of non-relevant documents. We broadly classify models as those that do not depend on other documents (Multi-Document-Agnostic, MDA) and those that do (Multi-Document-Dependent, MDD). Our results show that neither MDD nor MDA models are fully robust to the addition of non-relevant documents, as all models exhibit some performance degradation. Interestingly, among the models we test, MDA is more effective than MDD for retrieval, whereas MDD and MDA rerankers are equally effective.
CommentsCIKM 2026 Short Paper track