发表机构
Sun Yat-sen University; Central South University(中山大学; 中南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
HC-RAG是面向异构金融文件的以证据为中心的分层跨模态RAG框架,通过构建类型化金融证据图、按意图路由证据,在金融问答基准上较RAPTOR、GraphRAG取得显著性能提升。
AI 中文摘要
针对年度报告的金融问答不仅需要检索语义相似的段落,还常常涉及识别相关公司和财年、定位标准化文件章节、收集文本和表格证据,以及对照原始文档核查答案。然而,现有检索增强生成(RAG)系统通常将长文件展平为无序块,对金融报告的类型化结构关注有限,且采用固定的文本-表格融合策略,未考虑查询意图。为解决这些局限,我们提出HC-RAG,一种面向以证据为中心的金融问答的分层跨模态检索增强生成框架。HC-RAG将文件组织为包含文档、章节、文本单元、表格单元和元数据节点的类型化金融证据图,通过文档-章节-单元路径检索证据,在共享检索空间中对齐文本与表格证据,并根据计算、趋势、事实、比较四种语义意图路由证据。我们进一步推出Multi-Doc-2025基准,包含来自87家标普500公司179份美国证券交易委员会(SEC)10-K文件的2327个经专家验证的问答对,带有意图、难度和结构化证据属性标签。在公开金融问答基准及Multi-Doc-2025上的实验表明,HC-RAG提升了答案质量和证据定位能力,尤其在长文档、表格相关及跨文档场景中表现突出;在DocFinQA上,HC-RAG较RAPTOR的F1值提升6.6个百分点,在Multi-Doc-2025上较GraphRAG的F1值提升10.9个百分点。证据级分析与 ablation 研究显示,改进主要源于更准确的章节定位、表格锚定、跨文档证据聚合及意图感知的文本-表格路由。
英文摘要
Financial question answering over annual reports requires more than retrieving semantically similar passages. It often involves identifying relevant companies and fiscal years, locating standardized filing sections, collecting textual and tabular evidence, and checking answers against the original documents. Existing RAG systems, however, usually flatten long filings into unordered chunks, pay limited attention to the typed structure of financial reports, and use fixed text-table fusion strategies without considering query intent. To address these limitations, we propose \textbf{HC-RAG}, a hierarchical cross-modal retrieval-augmented generation framework for evidence-centric financial QA. HC-RAG organizes filings into a typed financial evidence graph with documents, sections, text units, table units, and metadata nodes. It retrieves evidence through document-section-unit paths, aligns textual and tabular evidence in a shared retrieval space, and routes evidence according to four semantic intents: calculation, trend, fact, and comparison. We further introduce \textbf{Multi-Doc-2025}, a benchmark containing 2,327 expert-verified QA pairs from 179 SEC 10-K filings of 87 S\&P 500 companies across fiscal years 2022--2024, with labels for intent, difficulty, and structural evidence attributes. Experiments on public financial QA benchmarks and Multi-Doc-2025 show that HC-RAG improves both answer quality and evidence localization, especially in long-document, table-related, and cross-document settings. HC-RAG outperforms RAPTOR by 6.6 F1 points on DocFinQA and GraphRAG by 10.9 F1 points on Multi-Doc-2025. Evidence-level analysis and ablation studies show that the improvements mainly come from more accurate section localization, table grounding, cross-document evidence aggregation, and intent-aware text-table routing.
Comments16 pages, 5 figures