arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08940cs.IR

面向时序检索的推理视图难度门控融合

Difficulty-Gated Fusion of Reasoning Views for Temporal Retrieval

Jamie Holdcroft, Abdelrahman Abdallah, Adam Jatowt

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对时序检索中多改写视图融合的精度问题,提出查询难度门控的推理视图融合方法,在Tempo基准上提升了6种检索器的性能,且增益显著。

中文摘要 AI 辅助

依赖推理的时序检索需要将查询与那些相关性取决于共享时序推理而非词汇重叠的文档进行匹配。将查询扩展为多个明确表达其时序意图的改写版本,并分别进行检索,可提供所需的推理支持,但将得到的排名以相等权重融合会降低准确性:对于任意单个查询,仅部分改写版本可靠。我们提出查询难度门控的推理视图融合方法。从每个视图中读取其得分分布的8维特征,该特征由软max熵、得分间隙和离散度等查询性能预测量构建,一个约含1000个参数的门控网络将这些特征映射为每个查询的视图权重。融合排名在推理时无需相关性标签、重新排序或对检索器进行微调;门控网络采用留一任务交叉验证训练。在Tempo基准测试中,该方法提升了我们评估的全部6种检索器,从BERT编码器到7B解码器检索器,在较弱的骨干模型上提升幅度最大。最强的检索器达到0.297和0.303的nDCG@10,且与原始查询相比,每个查询的增益在配对自举检验下显著(p<0.001)。每个查询的最优组合达到0.364,而我们的实现结果为0.297,这一差距表明,针对每个查询的视图选择是未来的具体研究方向。

英文摘要

Reasoning-intensive temporal retrieval requires matching a query to documents whose relevance depends on shared temporal reasoning rather than lexical overlap. Expanding a query into several reformulations that make its temporal intent explicit, and retrieving with each, supplies this reasoning, but fusing the resulting rankings with equal weights wastes accuracy: for any single query, only some reformulations are reliable. We propose query-difficulty-gated fusion of reasoning views. From each view we read an eight-dimensional signature of its score distribution, built from query-performance-prediction quantities such as softmax entropy, score gaps, and dispersion, and a gate of roughly one thousand parameters maps these signatures to per-query view weights. The fused ranking uses no relevance labels at inference, no re-ranking, and no fine-tuning of the retriever; the gate is trained leave-one-task-out. On the \textsc{Tempo} benchmark, the method improves all six retrievers we evaluate, from BERT encoders to 7B decoder retrievers, with the largest gains on the weaker backbones. The strongest retrievers reach $0.297$ and $0.303$ nDCG@10, and the per-query gain over the original query is significant under a paired bootstrap ($p<0.001$). A per-query oracle reaches $0.364$ against our realized $0.297$, exposing headroom that identifies per-query view selection as a concrete next step.

补充信息

↑