arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FinSkillBench:评估用于投资管理的智能体AI与领域技能

FinSkillBench: Evaluating AI Agents and Domain Skills for Investment Management

Jermyn Zhen Yong Bek, Zhuang Qiang Bok, Zhongtian Sun

arXiv 2608.18099首次发表:更新:

发表机构

Deep Insight Labs; University of Kent; University of Cambridge(深洞察实验室; 肯特大学; 剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FinSkillBench是评估投资管理领域AI智能体金融技能的基准,经9个模型的大规模评估发现,精选技能可显著提升性能,而自生成技能益处有限,其成果为该领域研究提供了支持。

AI 中文摘要

投资管理是一个高风险领域,其中具智能体的AI系统必须做的不仅仅是生成看似合理的文本。它们必须检索时点数据、组装正确的计算输入、调用专门方法,并生成可审计的结构化输出。我们推出FinSkillBench,这是一个旨在衡量语言模型智能体能否有效运用金融领域技能解决投资管理任务的评估套件。该基准涵盖投资组合构建、风险管理和基本面分析三个领域,包含12个子任务,共2603个任务片段。每个片段提供时点输入、隐藏的真实值以及特定于任务的链接。我们比较三种条件:无技能、由程序文档和可执行组件组成的精选技能包,以及智能体在一个片段内编写并复用自身程序的自生成技能。在9个模型的大规模评估中,精选技能持续提升性能,将平均得分从0.366提升至0.528,在投资组合构建和风险管理中提升最大。相比之下,自生成技能尽管计算成本更高,但几乎没有带来益处。使用独立智能体框架Hermes Agent(8个模型,总计5280个片段)进行的独立评估,在所有三个领域都重现了这一方向模式,技能效果的幅度随子任务和工具的不同而变化。这些结果表明,在投资管理智能体中,获取可靠的程序技能与模型选择同等重要,而单纯的技能自生成往往无效。我们发布该基准、评估工具、精选技能包和完整轨迹,以支持进一步研究。

英文摘要

Investment management is a high-stakes domain in which agentic AI systems must do more than generate plausible text. They must retrieve point-in-time data, assemble correct computational inputs, invoke specialized methods, and produce auditable structured outputs. We introduce FinSkillBench, an evaluation suite designed to measure whether language model agents can effectively use financial domain skills to solve investment management tasks. The benchmark spans three domains, portfolio construction, risk management, and fundamental analysis, and includes 12 subtasks with 2,603 task episodes. Each episode provides point-in-time inputs, hidden ground truth, and a task-specific verifier.We compare three conditions: no skill, curated skill packages consisting of procedural documents and executable components, and self-generated skills in which the agent writes and reuses its own procedures within an episode. Across 9 models and a large-scale evaluation, curated skills consistently improve performance, raising mean scores from 0.366 to 0.528, with the largest gains in portfolio construction and risk management. In contrast, self-generated skills provide little benefit despite higher computational cost. An independent evaluation using a separate agent framework (Hermes Agent, 8 models, 5,280 episodes total) reproduces the directional pattern across all three domains, with the magnitude of skill effects varying by subtask and harness. These results showthat in investment management agents, access to reliable procedural skills can be as important as model choice, while naive self-generation of skills is often ineffective. We release the benchmark, evaluation tools, curated skill packages, and full trajectories to support further research.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑