发表机构
Duke University(杜克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对金融文档检索问题,在修正语料库后发现简单融合BM25与紧凑密集编码器为强基线,而三种轻量自适应查询加权路由均未取得统计可靠提升。
AI 中文摘要
金融文件检索难度大,因为查询简短且包含大量缩写,而承载答案的证据位于冗长、表格密集的文档中。我们在FinDER(针对公司10-K文件的专家标注问题基准)上研究稀疏-密集混合检索。我们的第一个方法学发现是:若检索单元大于密集编码器的输入窗口,密集模型将无法看到大部分标注证据,从而混淆与全文稀疏基线的对比。我们直接测量该问题,并通过将语料库分割为编码器大小的窗口来消除它。在修正后的语料库上,融合BM25与紧凑密集编码器使参考级Hit@10较任一组件提升约28%,且在探索性对比中,无需训练、未调优的倒数秩融合优于等权重融合。随后我们探究按查询选择融合权重是否有帮助:插值权重网格的预言机显示有21.8%的提升空间,但三种轻量自适应路由(得分置信度启发式、基于查询特征的随机森林、基于查询嵌入的岭回归)在按公司分组的交叉验证及聚类稳健推断下,均未较固定融合建立统计可靠的提升。简单融合在此是强基线,我们讨论了按查询加权无法利用可用提升空间的原因。
英文摘要
Retrieval over financial filings is difficult because queries are short and acronym-heavy while the answer-bearing evidence sits inside long, table-dense documents. We study sparse-dense hybrid retrieval on FinDER, a benchmark of expert-annotated questions over corporate 10-K filings. Our first finding is methodological: if the retrieval unit is larger than the dense encoder's input window, the dense model never sees a large share of the labeled evidence, confounding comparison against a full-text sparse baseline. We measure this directly and reduce the imbalance with windows chosen for the encoder budget. On this corpus, fusing BM25 and a compact dense encoder improves reference-level Hit@10 by roughly 28 percent over either component, and training-free, untuned reciprocal rank fusion exceeds the equal-weight blend in an exploratory comparison. We then ask whether choosing the fusion weight per query helps: an oracle over the interpolation-weight grid shows headroom of 21.8 percent, yet none of the three lightweight adaptive routers (a score-confidence heuristic, a random forest over query features, and a ridge regressor over query embeddings) establishes a statistically reliable improvement over the fixed blend under company-grouped cross-validation with cluster-robust inference. Simple fusion is a strong baseline here, and we discuss why per-query weighting does not capture the available headroom.