发表机构
University of Groningen; Goethe University Frankfurt; DataNXT GmbH; Halle Institute for Economic Research (IWH); Martin Luther University Halle-Wittenberg(格罗宁根大学; 法兰克福大学; DataNXT有限公司; 哈雷经济研究所; 马丁路德大学哈雷-维滕贝格分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FinRank是针对SEC文件金融问答与检索的基准,含1185条人工问答记录,评估段落检索等任务,实验显示现有模型性能有限,为开发可靠金融问答系统提供支撑。
AI 中文摘要
金融问答通常以答案正确性进行评估,但在SEC文件中,看似合理甚至数值正确的答案可能基于错误的证据。类似事实和披露会在同一文件的不同章节、同一公司的不同报告期以及可比公司之间重复出现。FinRank针对这种对证据来源敏感的检索问题,要求系统识别出针对目标实体、报告期和披露背景的证据。该基准包含22家公司的10-K和10-Q文件上的1185个人工撰写的问答记录,每条记录包含参考答案、黄金支持段落,以及从文件内易混淆段落、不同报告期和可比公司中精心筛选的难负例。FinRank将段落检索、重排序和难负例判别作为单独任务进行评估。基线结果表明该设置的难度:在评估的系统中,即使是7B指令调优嵌入器在合并证据语料库上的Recall@10仅达44.8%;亚十亿参数编码器较BM25最多提升3.5个百分点,金融适配嵌入器较BM25落后9.7个百分点,当用精心筛选的难负例替换随机负例时,成对准确率下降13.0-20.5个百分点。FinRank为开发不仅准确且基于正确披露的金融问答系统提供了以证据为先的基准。
英文摘要
Financial question answering is typically evaluated by answer correctness, yet in SEC filings a plausible and even numerically correct answer can be grounded in the wrong evidence. Similar facts and disclosures recur across sections of a filing, across reporting periods of the same firm, and across comparable firms. FinRank targets this provenance-sensitive retrieval problem by requiring systems to identify evidence for the intended entity, reporting period, and disclosure context. The benchmark contains 1185 manually authored question-answer records over the 10-K and 10-Q filings of 22 companies. Each record includes a reference answer, gold supporting passages, and hand-curated hard negatives drawn from confusable passages within filings, across reporting periods, and across comparable firms. FinRank evaluates passage retrieval, reranking, and hard-negative discrimination as separately measured tasks. Baseline results demonstrate the difficulty of this setting: among the evaluated systems, even a 7B instruction-tuned embedder reaches only 44.8% Recall@10 on the pooled evidence corpus; sub-billion-parameter encoders gain at most 3.5 points over BM25, a finance-adapted embedder trails BM25 by 9.7 points, and pairwise accuracy falls by 13.0-20.5 percentage points when random negatives are replaced with the curated hard negatives. FinRank provides an evidence-first benchmark for developing financial question answering systems that are not only accurate but also grounded in the correct disclosure.
Comments24 pages, 3 figures. Dataset and evaluation code: https://github.com/datanxt/FinRank