arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03564cs.AI

知识还是计算器?可验证金融智能体工作流中的技能溢价分解

Knowledge or Calculator? Decomposing the Skill Premium in Verifiable Financial Agent Workflows

  • Deep Insight Labs(深度洞察实验室)
  • University of Kent(肯特大学)
  • University of Cambridge(剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

Jermyn Zhen Yong Bek, Zhuang Qiang Bok, Zhongtian Sun

AI总结:

提出FinSkillBench基准,分解金融智能体工作流中的技能溢价,发现可执行工具贡献最大,文档次之,组合呈次可加性,溢价依赖系统整体。

AI中文摘要:

金融AI智能体必须做的不仅仅是检索事实:投资工作流需要正确的定量执行、对程序性资源的可靠使用,以及可审计的结构化输出。我们引入了FinSkillBench,一个包含投资组合构建、风险管理和基本面分析中12个子任务的2,603个时点情景的评估套件,具有隐藏的可再生成地面真值和任务特定的确定性验证器。执行了9个模型和3种资源条件下的17,820个情景,对8个模型的配对分析表明,精心策划的技能包将平均得分提高了+16.2分(从0.366到0.528),而单个情景内生成的技能仅增加+0.5分,同时消耗更多的令牌和轮次。然后,我们通过分别授予人工编写的程序性文档和可执行的领域工具来分解策划溢价:仅文档增加+5.6分,仅工具增加+19.5分,而它们的组合是次可加的。该溢价强烈依赖于工作流:可执行工具在数值密集型工作流中占主导地位,当程序性或输出模式指导成为瓶颈时,文档更为重要,而解释性任务则两者兼受益。这些效应在10种评分变体和聚类自助分析中符号稳定,独立实施的第二测试平台重现了方向性模式,同时表明效应大小取决于工具和数据如何暴露。总体而言,测得的“技能溢价”是整个模型、资源和测试平台系统的属性,而非仅底层模型本身的属性。

英文摘要:

Financial AI agents must do more than retrieve facts: investment workflows require correct quantitative execution, reliable use of procedural resources, and auditable structured outputs. We introduce FinSkillBench, an evaluation suite of 2,603 point in time episodes across 12 subtasks in portfolio construction, risk management, and fundamental analysis, with hidden regenerable ground truth and task specific deterministic verifiers. Executing 17,820 episodes across 9 models and 3 resource conditions, the paired analysis across 8 models shows that curated skill packages raise mean scores by +16.2 points (0.366 to 0.528), whereas skills generated within a single episode add only +0.5 points while consuming more tokens and turns. We then decompose the curated premium by granting human authored procedural documents and executable domain tools separately: documents alone add +5.6 points, tools alone add +19.5 points, and their combination is subadditive. The premium is strongly workflow dependent: executable tools dominate numerically intensive workflows, documentation matters more when procedural or output schema guidance is the bottleneck, and interpretive tasks benefit from both. The effects are sign stable across 10 scoring variants and cluster bootstrap analyses, and an independently implemented second harness reproduces the directional pattern while showing that effect magnitudes depend on how tools and data are exposed. Overall, a measured "skill premium" is a property of the full model, resource, and harness system rather than of the underlying model alone.

补充信息

↑