arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11683cs.AIcs.CL

FrontierFinance:用于衡量金融智能体前沿智能的具有挑战性的基准

FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents

  • Samaya AI(萨马亚人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

Yuhao Zhang, O. Ozan Koyluoglu, Thejas Venkatesh, Richard Diehl Martinez, Vishank Bhatia, Arash Alidoust, Ashwin Paranjape

AI总结:

该研究推出覆盖投资者全流程的FrontierFinance基准,评估发现工具框架影响智能体表现,Samaya内部系统最优,开源模型Kimi K3成本更低且接近专有模型,最难用例为筛选与发现等。

AI中文摘要:

人工智能智能体越来越多地被部署用于专业投资研究,但目前尚无基准能够覆盖完整投资者工作流程的复杂性。现有基准主要聚焦于金融数据提取这一狭窄领域,当前模型在该领域已基本达到饱和状态,而基于参考的指标以及通用大语言模型(LLM)作为评判者的评分方式,无法满足真实分析师查询所需的开放式、长篇答案的评估需求。我们推出了FrontierFinance,这是一个完全开源的基准,包含220个由专家精心设计的查询,以及11543个带有来源标注的评分规则,覆盖了完整投资者工作流程中的六个关键用例。FrontierFinance比现有的公开金融基准更广泛、更具挑战性。在仅使用公开可用数据的统一工具框架下评估前沿模型和智能体系统时,我们发现工具框架而非模型本身对质量和效率有显著影响;Samaya的内部系统以56.0%的表现领先,其成本约为最强前沿模型Claude Fable 5(49.2%)的2.2倍;性能最佳的开源权重模型Kimi K3(46.4%)的表现接近最佳专有模型,成本仅为其4.5倍。在所有系统中,筛选与发现、行业及宏观分析仍是最具挑战性的用例,即便表现最佳的系统也仅达到33%和39%。我们将公开该数据集和评分代码。

英文摘要:

AI agents are increasingly deployed for professional investment research, yet no benchmark captures the complexity of the full investor workflow. Existing benchmarks mainly target financial data extraction, a narrow slice that current models have largely saturated, while reference-based metrics and generic LLM-as-a-judge scoring fall short on the open-ended, long-form answers that real analyst queries demand. We introduce FrontierFinance, a fully open benchmark of 220 expert-crafted queries and 11,543 source-attributed rubrics spanning six crucial use cases across the full investor workflow. FrontierFinance is both broader and harder than existing public finance benchmarks. Evaluating frontier models and agent systems under a common harness restricted to publicly available data, we find that the tool harness, not the model alone, strongly shapes quality and efficiency; that Samaya's in-house system leads at 56.0%, ahead of the strongest frontier model (Claude Fable 5, 49.2%) at roughly 2.2x lower cost; and that the best open-weight model (Kimi K3, 46.4%) nearly matches the best proprietary model at 4.5x lower cost. Screening & Discovery and Sector, Industry & Macro remain the hardest use cases across all systems, where even the best systems reach only 33% and 39%. We make the dataset and grading code publicly available.

↑