AI 中文总结
研究针对金融文档智能体推理性能差异问题,设计Finance-LaTeX SKILL生成金融文档和问答对,引入FinanceComplexQA基准测试,对领先系统和工具全面评估,通过分析失败案例研究其在多方面的能力。
AI 中文摘要
智能体推理因其整合大规模信息并生成可靠准确内容的能力,已成为金融分析中的变革力量。但处理复杂现实问题时,不同智能体性能仍有显著差异。本文设计Finance-LaTeX SKILL,用于合成复杂布局金融文档。基于此生成2000份专业金融文档及6000个高质量问答对。引入FinanceComplexQA基准测试,含2026个针对1009份金融文档的深度研究任务,有双语支持等8个关键特性。利用其对领先的RAG系统和智能体推理工具进行全面评估,通过分析失败案例深入研究其在数值计算等方面的能力。
英文摘要
Agentic Reasoning has become a transformative force in financial analysis due to its ability to integrate large-scale information and generate reliable and accurate content. However, when handling complex real-world problems, different agents still show significant performance variation. In this work, we design Finance-LaTeX SKILL, a skill for synthesizing financial documents with complex layouts based on expert knowledge. Using an agent workflow built on this skill, we generate 2,000 professional financial documents along with 6,000 high-quality question-answer pairs. To evaluate the overall capability of agents, we introduce FinanceComplexQA, a comprehensive open-ended generation benchmark for financial documents that closely resembles real-world scenarios. It contains 2,026 deep research tasks targeting 1009 financial documents. FinanceComplexQA has 8 key features: bilingual support; coverage of six mainstream scenarios and seven tasks; expert-level document reasoning questions; deep research of complex layouts; relatively stable and permanent reference answers; and precise evaluation through an Agent-as-a-Judge with multiple evaluation metrics. Using FinanceComplexQA, we conduct a comprehensive evaluation of leading RAG systems and agentic reasoning tools for financial document QA. Through identifying and analyzing failure cases, we provide an in-depth study of their capabilities in numerical computation, multi-hop reasoning, content summarization, and industry analysis.
Comments27 pages, 9 tables, 2 figures