TF-IDF 和 BM25 是精确的 KL 散度
TF-IDF and BM25 Are Exact KL Divergences
浏览论文内容
中文总结 AI 辅助
本文证明 TF-IDF 和 BM25 可精确解释为两个概率模型间的 KL 散度,为两者提供统一理论框架,并支持与其他检索方法的理论比较。
中文摘要 AI 辅助
TF-IDF 和 BM25 是用于评分查询-文档相关性最广泛使用的两种方法,然而两者都没有标准的概率推导,以证明它们在统一框架内作为统计方法的合理性。我们通过证明这两种评分方法都允许被精确解释为两个概率模型之间的 Kullback-Leibler 散度来弥补这一空白。我们处理包含 IDF 项中加 1 修正的 BM25 变体,这是实践中使用的版本,同时也讨论了没有该修正的原始 BM25 公式。由此产生的框架为 TF-IDF 和 BM25 提供了共同的理论基础,阐明了它们衡量的内容,并允许它们与其他信息检索方法进行理论比较,而不仅仅是实验比较。
英文摘要
TF-IDF and BM25 are two of the most widely used methods for scoring query-document relevance, yet neither has a standard probabilistic derivation that justifies it as a statistical method within a unified framework. We address this gap by showing that both scoring methods admit an exact interpretation as Kullback-Leibler divergences between two probability models. We treat the BM25 variant that includes the plus 1 correction in the IDF term, which is the one used in practice, and also discuss the original BM25 formulation without that correction. The resulting framework provides a common theoretical basis for TF-IDF and BM25, clarifies what they measure, and allows them to be compared theoretically with other information retrieval methods rather than only experimentally.