arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于锚点的点式大语言模型重排序器何时有用?检索器质量、统计范围与锚点设计

When Do Anchor-Based Pointwise LLM Rerankers Help? Retriever Quality, Statistical Scope, and Anchor Design

Utshab Kumar Ghosh, Shubham Chatterjee

arXiv 2608.10528首次发表:更新:

发表机构

Missouri University of Science and Technology(密苏里科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究以GCCP/PAGC为对象,复现并分析锚点式点式LLM重排序器,发现其增益源于对比打分,适用条件较窄,锚点构建可简化,检索器强度影响其效果。

AI 中文摘要

基于锚点的点式大语言模型(LLM)重排序器会针对每个候选段落与共享参考段落打分,以在点式成本下恢复跨文档上下文。我们以GCCP/PAGC为代表性方法,研究其真正发挥作用的场景,研究以复现为首要目标。我们将复现作为起点,对基于锚点的点式重排序开展受控的组件级压力测试。仅基于论文文本的初始复现实现了0.24的nDCG@10,远低于报告的0.66,这表明要复现该方法需多个未记录的实现细节。在识别并补全8个此类细节后,我们复现了报告结果,误差在1.6%以内,并使用经验证的实现开展受控分析。我们发现核心对比打分思想在严格统计修正下具有鲁棒性,但论文中固定的两项设计选择可靠性较低:其一,将对比打分与标准点式相关性打分结合,在第一阶段检索器为BM25时有助益,但在第一阶段检索器为更强的密集模型(如E5)时几乎无增益;其二,论文中更复杂的锚点构建方法并非必要,通过交错排名最高的句子构建的简单锚点,在各数据集上的表现与复杂锚点相当或更优。这些发现在不同LLM骨干(包括4位量化的72B模型)上一致。总体而言,基于锚点的点式重排序是有效的,但其增益主要来自对比打分,而非更复杂的聚合与锚点构建选择,且其适用条件比原始评估显示的更狭窄。

英文摘要

Anchor-based pointwise LLM reranking scores each candidate against a shared reference passage to recover cross-document context at pointwise cost. We study when this actually helps, using GCCP/PAGC as a representative method. Our study is reproduction-first. We use reproduction as a starting point for a controlled component-level stress test of anchor-based pointwise reranking. Our initial reimplementation, based only on the paper text, achieves 0.24 nDCG@10 instead of the reported 0.66, revealing that several undocumented implementation details are necessary to reproduce the method. After identifying and recovering eight such details, we reproduce the reported results within 1.6% and use the validated implementation for controlled analysis. We find that the core contrastive scoring idea is robust under rigorous statistical correction. However, two design choices held fixed in the original paper are less reliable. First, we find that combining the contrastive score with the standard pointwise relevance score helps when the first-stage retriever is BM25, but gives little or no benefit when the first-stage retriever is a stronger dense model such as E5. Second, the paper's more complex method for constructing the anchor is unnecessary. A much simpler anchor, built by interleaving the top-ranked sentences, matches or outperforms it across datasets. These findings are consistent across different LLM backbones, including a 4-bit quantized 72B model. Overall, anchor-based pointwise reranking is effective, but its gains come mainly from contrastive scoring rather than from the more complex aggregation and anchor-construction choices, and they appear under narrower conditions than the original evaluation suggests.

CommentsTo be published in the 35th ACM International Conference on Information and Knowledge Management (CIKM 2026)

DOI:10.1145/3799682.3841055

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑