FinFIRST:面向金融信息检索、溯源与可追踪性的搜索智能体基准
FinFIRST: Benchmarking Search Agents for Financial Information Retrieval, Sourcing and Traceability
浏览论文内容
中文总结 AI 辅助
FinFIRST是首个联合评估金融搜索智能体答案与证据的基准,通过原子化评分和123个专家任务,使研究过程可测量、可验证、可诊断。
中文摘要 AI 辅助
金融搜索对LLM智能体而言是一项要求极高的任务,不仅需要给出正确的最终答案,还需要具备时间有效的检索、权威来源选择、实体与时期对齐、单位与定义一致性,以及所有结论的可验证证据。现有基准主要仅评估最终答案,这使得难以定位错误或评估答案是否有充分依据。为弥补这一空白,我们提出了FinFIRST(金融信息检索、溯源与可追踪性),这是首个通过原子化评分细则同时评估答案与支持证据的金融基准。FinFIRST包含123个由专家撰写的任务,覆盖从易到难的梯度难度谱系,这些任务基于真实金融场景的聚合模式构建,通过一个18字段分类体系、一个六轴覆盖蓝图、一个包含138个金融来源的注册表、超过50位金融专家的贡献以及一个六阶段质量控制流程来生成。每个任务附带一个基于证据的参考包,该包被分解为三个维度的原子化标准:原始信息获取、来源验证、以及计算与答案形成。我们在统一工具设置下评估了15种模型配置。Claude-Opus-5取得了最高的原子化得分87.59%,而GPT-5.6-Sol达到了最高的严格通过率71.54%。在所有系统中,计算与答案形成环节持续落后于原始信息获取环节。FinFIRST将最终答案正确性作为首要目标,同时使支持性研究过程变得可测量、可验证和可诊断。
英文摘要
Financial search is a highly demanding task for LLM agents, requiring not only a correct final answer but also temporally valid information retrieval, authoritative source selection, entity and period alignment, unit and definition consistency, and verifiable evidence for all conclusions. Existing benchmarks predominantly evaluate only the final answer, making it difficult to localize errors or assess whether an answer is well-founded. To address this gap, we introduce FinFIRST (Financial Information Retrieval, Sourcing and Traceability), the first financial benchmark to jointly evaluate answers and supporting evidence through atomic rubrics. FinFIRST comprises 123 expert-authored tasks spanning a graduated difficulty spectrum, constructed from aggregate patterns of real-world financial scenarios through an 18-field taxonomy, a six-axis coverage blueprint, a registry of 138 financial sources, contributions from over 50 finance experts, and a six-stage quality-control pipeline. Each task is accompanied by an evidence-grounded reference package decomposed into atomic criteria across three dimensions: raw-information acquisition, source verification, and computation and answer formation. We evaluate 15 model configurations under a unified tool setting. Claude-Opus-5 achieves the highest atomic score of 87.59%, while GPT-5.6-Sol attains the highest strict pass rate of 71.54%. Computation and answer formation consistently lag behind raw-information acquisition across systems. FinFIRST retains final-answer correctness as the primary objective while making the supporting research process measurable, verifiable, and diagnosable.
发表机构
- Ling Team, Inclusion AI(灵团队,Inclusion AI)
- Peking University(北京大学)
- China International Capital Corporation Limited(中国国际金融股份有限公司)
机构由 AI 辅助整理,请以论文原文为准。