增强金融问答:一种新的银行财务报表基准数据集
Enhancing Financial Question Answering: A Novel Benchmark Dataset of Banks' financial statements
浏览论文内容
中文总结 AI 辅助
针对银行财务报表对比分析的挑战,本文构建了新型金融问答基准数据集FinRAG-QA,评估多阶段RAG流水线各组件,提升了检索与问答性能。
中文摘要 AI 辅助
银行财务报表的对比分析对自动问答系统构成重大挑战,因其具有复杂性、篇幅长、专业术语多,且不同司法管辖区和机构的文本与数值内容存在异质性。本文推出FinRAG-QA,一种用于金融问答的新型基准数据集,包含由从业者策划的999个问题,涉及10项标准化指标,基于24家欧美主要银行2019至2023年间的209份年度报告和第三支柱报告构建。与此前聚焦美国文件及单一机构分析的金融问答基准不同,FinRAG-QA针对跨机构检索,其文档平均长度达19.8万字,长于现有任何金融问答资源。我们在该基准上评估了多阶段RAG流水线,并分离各组件的贡献:上下文块增强结合检索优化型嵌入模型,使NDCG@10从0.322提升至0.710;在检索到真实答案的前提下,推理优化型生成器将答案准确率从44.6%提升至79.0%(提升34.4个百分点),生成延迟约为原来的20倍。我们进一步发现,当第一阶段排序已足够强时,交叉编码器重排序会降低检索性能,且在生成时,单个排名第一的块优于更大的上下文。实验于2024年末至2025年初开展,使用了当时可用的模型。
英文摘要
The comparative analysis of banks' financial statements poses significant challenges for automated question answering systems due to their complexity, substantial length, technical language, and inhomogeneity of both textual and numerical content across different jurisdictions and institutions. We introduce FinRAG-QA, a novel benchmark dataset for financial question answering, which comprises 999 practitioner-curated questions on 10 standardised indicators, grounded in 209 annual and Pillar 3 reports from 24 major European and U.S. banks spanning 2019-2023. Unlike prior financial QA benchmarks, which centre on U.S. filings and single-institution analysis, FinRAG-QA targets cross-institutional retrieval over documents averaging 198k words, longer than any existing financial QA resource. On this benchmark we evaluate a multi-stage RAG pipeline and isolate the contribution of each component. Contextual chunk enrichment combined with a retrieval-optimised embedding model raises NDCG@10 from 0.322 to 0.710; conditional on the ground truth being retrieved, a reasoning-optimised generator raises answer accuracy from 44.6% to 79.0% (+34.4 percentage points), at roughly 20x the generation latency. We further show that cross-encoder reranking degrades retrieval when the first-stage ranking is already strong, and that a single top-ranked chunk outperforms larger contexts at generation time. Experiments were run in late 2024-early 2025 with the models available at that time.