arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FinanceComplexQA:对工业级金融文档进行智能体推理的基准测试

FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents

Xianfu Cheng, Shiwei Zhang, Jiyu Zhao, Jian Yang, Xinyuan Wang, Ming Zhou, Weixiao Zhou, Xiangyuan Guan, Xiang Li, Zhenhe Wu, Ziyi Ni, Zhoujun Li, Bingjing Xu

arXiv 2607.19238首次发表:更新:

AI 中文总结

研究针对金融文档智能体推理性能差异问题,设计Finance-LaTeX SKILL生成金融文档和问答对,引入FinanceComplexQA基准测试,对领先系统和工具全面评估,通过分析失败案例研究其在多方面的能力。

AI 中文摘要

智能体推理因其整合大规模信息并生成可靠准确内容的能力,已成为金融分析中的变革力量。但处理复杂现实问题时,不同智能体性能仍有显著差异。本文设计Finance-LaTeX SKILL,用于合成复杂布局金融文档。基于此生成2000份专业金融文档及6000个高质量问答对。引入FinanceComplexQA基准测试,含2026个针对1009份金融文档的深度研究任务,有双语支持等8个关键特性。利用其对领先的RAG系统和智能体推理工具进行全面评估,通过分析失败案例深入研究其在数值计算等方面的能力。

英文摘要

Agentic Reasoning has become a transformative force in financial analysis due to its ability to integrate large-scale information and generate reliable and accurate content. However, when handling complex real-world problems, different agents still show significant performance variation. In this work, we design Finance-LaTeX SKILL, a skill for synthesizing financial documents with complex layouts based on expert knowledge. Using an agent workflow built on this skill, we generate 2,000 professional financial documents along with 6,000 high-quality question-answer pairs. To evaluate the overall capability of agents, we introduce FinanceComplexQA, a comprehensive open-ended generation benchmark for financial documents that closely resembles real-world scenarios. It contains 2,026 deep research tasks targeting 1009 financial documents. FinanceComplexQA has 8 key features: bilingual support; coverage of six mainstream scenarios and seven tasks; expert-level document reasoning questions; deep research of complex layouts; relatively stable and permanent reference answers; and precise evaluation through an Agent-as-a-Judge with multiple evaluation metrics. Using FinanceComplexQA, we conduct a comprehensive evaluation of leading RAG systems and agentic reasoning tools for financial document QA. Through identifying and analyzing failure cases, we provide an in-depth study of their capabilities in numerical computation, multi-hop reasoning, content summarization, and industry analysis.

Comments27 pages, 9 tables, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑