arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FinanceHarness:自主金融深度研究框架

FinanceHarness: Autonomous Financial Deep Research Framework

Yijia Xiao, Rujun Han, Yanfei Chen, Zifeng Wang, Ke Jiang, Zhongying CuiZhu, Vishy Tirumalashetty, Wei Wang, Burak Gokturk, Tomas Pfister, Chen-Yu Lee

arXiv 2607.27853首次发表:更新:

AI 中文总结

FinanceHarness是一款端到端自动化金融深度研究的自主框架,结合自主智能体与金融专用工具,在FinanceGym基准上大幅提升了金融研究评分规则得分。

AI 中文摘要

得益于大型语言模型(LLMs)和自主智能体的进展,深度研究已成为应用最广泛的智能体产品之一。然而,大多数深度研究系统生成通用报告,无法满足金融深度研究的需求,金融研究需要专业知识来分析历史模式并预测未来事件。因此,自动化金融深度研究既需要分层框架来驱动研究智能体,也需要可验证的时点基准以防止未来信息泄露。我们提出FinanceHarness,该框架运行面向金融的工具和从业者指导的工作流,端到端自动化金融深度研究:环境与数据构建、智能体执行循环以及奖励建模。我们进一步提出FinanceGym,包含由论文驱动的研究问题和结合截止前与截止后标准的评分规则。专业专家验证显示其通过率为82%,即使是领先的LLMs和智能体在该评分规则上的得分也低于40%,表明FinanceGym具有挑战性且存在大量提升空间。使用相同的开源权重主干模型,FinanceHarness将整体评分规则得分从25.3%提升至32.4%。FinanceHarness可通过此URL获取。

英文摘要

Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research systems write general-purpose reports, which are inadequate for financial deep research. Financial research demands specialized knowledge to analyze historical patterns and forecast upcoming events. Automating financial deep research therefore requires both a layered harness to drive the research agent and a verifiable, point-in-time benchmark that prevents leakage of future information. We present FinanceHarness, a harness that runs finance-oriented tools and practitioner-guided workflows, automating financial deep research end to end: environment and data construction, the agent execution loop, and reward modeling. We further propose FinanceGym, comprising thesis-driven research questions and rubrics that combine pre-cutoff and post-cutoff criteria. Professional expert validation yields an 82% pass rate. With the same open-weight backbone, FinanceHarness improves the overall rubric score from 25.3% to 32.4%, demonstrating the effectiveness of our specialized harness design. However, even pairing FinanceHarness with the most cutting edge LLM (e.g. Opus-5), the FinanceGym score is below 45%, showing that it is a challenging benchmark for financial deep research. Leaderboard is available at: https://financegym.github.io/ and FinanceHarness code is available at: https://github.com/Yijia-Xiao/FinanceHarness.

CommentsFinanceHarness available at https://github.com/Yijia-Xiao/FinanceHarness

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑