arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16386cs.CLcs.LG

Mint-Agent:引入金融原生智能体基础模型

Mint-Agent: Introducing Finance-Native Agentic Foundation Models

Agent Team, Kun Wang, Gavin Zhang, Yaze Geng, Lei Tang, Yaoyang Yi, Zonghan Wu, Yifan Hu, Qingsong Wen, Yilei Shao

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出金融原生智能体模型Mint-Agent,通过三大支柱开发出Mint-Cu(9B)和Mint-Ag(27B),在多个金融基准测试中展现出优异的可靠性与可执行性,为可信金融智能提供了新路径。

中文摘要 AI 辅助

金融智能体不仅需要回忆领域知识,还必须具备可靠性(基于扎实证据执行精确操作)和可执行性(维持结论可审计的长周期研究)。我们提出Mint-Agent,一系列围绕这两个金融智能规模设计的金融原生智能体模型。Mint-Agent构建于三大支柱:数据、harness和算法。我们的数据引擎从真实金融来源构建针对原子金融能力和长周期智能体执行的干净专用任务。MintHarness支持与开放环境的稳定交互,并在扩展研究轨迹中维护可审计的证据链。我们的训练配方结合SFT、关键步骤OPD和RLVR,开发独立的金融推理和智能体执行专家,随后通过模型合并和多教师在线策略蒸馏将其统一为紧凑通用金融智能体。该流程产生两个旗舰模型:Mint-Cu(9B)和Mint-Ag(27B)。在专业金融基准测试中,我们的模型展现两个核心优势:(1)可靠性:Mint-Ag在RFC-Bench上达到98.33%,超过GPT-5.6-Sol和Claude-Opus-4.8 3.66和3.00个百分点;(2)可执行性:Mint-Cu在FinSearchComp T2上达到69.86%,超过Agents-A1-35B和Nex-N2-mini 22.83和12.78个百分点,而Mint-Ag在FinanceAgentBench v1.1和v2上分别达到76.00%和60.49%。这些结果为构建可信金融智能指明了方向,其中领域专业知识、长周期执行和可审计证据被联合设计为前沿智能体模型的统一基础。

英文摘要

Financial agents must do more than recall domain knowledge: they must be both reliable, executing precise operations over grounded evidence, and executive, sustaining long-horizon research whose conclusions remain auditable. We present Mint-Agent, a family of finance-native agentic models designed around these two scales of financial intelligence. Mint-Agent is built upon three pillars: data, harness, and algorithm. Our data engine constructs clean, specialized tasks for atomic financial capabilities and long-horizon agentic execution from real-world financial sources. MintHarness enables stable interaction with open-ended environments and maintains auditable evidence trails across extended research trajectories. Our training recipe combines SFT, critical-step OPD, and RLVR to develop separate financial reasoning and agentic execution experts, which are then unified through model merging and multi-teacher on-policy distillation into compact, general-purpose financial agents. This pipeline yields two flagship models, Mint-Cu (9B) and Mint-Ag (27B). Across professional financial benchmarks, our models demonstrate two defining strengths: (1) Reliability: Mint-Ag achieves 98.33% on RFC-Bench, surpassing GPT-5.6-Sol and Claude-Opus-4.8 by 3.66 and 3.00 points; and (2) Executability: Mint-Cu reaches 69.86% on FinSearchComp T2, outperforming Agents-A1-35B and Nex-N2-mini by 22.83 and 12.78 points, while Mint-Ag achieves 76.00% and 60.49% on FinanceAgentBench v1.1 and v2, respectively. These results establish a path toward trustworthy financial intelligence in which domain expertise, long-horizon execution, and auditable evidence are jointly engineered as a unified foundation for frontier agentic models.

↑