发表机构
School of Computing, National University of Singapore; China Economics and Management Academy, Central University of Finance and Economics; Asian Institute of Digital Finance, National University of Singapore; Institute of Computing Technology, Chinese Academy of Sciences(新加坡国立大学计算机学院; 中央财经大学中国经济与管理研究院; 新加坡国立大学亚洲数字金融研究所; 中国科学院计算技术研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出首个评估深度研究智能体端到端金融指标构建表现的基准FinDeepIndicator,含多市场数据,实验发现DR智能体优于检索增强LLMs但仍不可靠,为金融领域智能体开发提供参考。
AI 中文摘要
金融指标是将原始金融数据转化为可解释度量的重要工具,适用于估值、风险评估、经济分析等各类下游任务。然而,现有金融基准大多聚焦于答案层面的准确性,且通常假设相关数据已提供,对指标构建中间过程的评估研究不足。本研究提出FinDeepIndicator,这是首个专门用于评估深度研究(Deep Research, DR)智能体在端到端金融指标构建任务中表现的基准。具体而言,FinDeepIndicator从指标构建的四个阶段对DR智能体进行评估:公式规范定义、数据收集、指标计算及答案生成,涵盖基本面、技术面和宏观经济指标,共分为21个细分子类别。该基准包含3350个精心整理的问答对,数据源自美国和中国市场,还包含10年的历史金融数据及800家上市公司的相关信息。对配备搜索功能的大型语言模型(Large Language Models, LLMs)和DR智能体开展的大量实验显示,LLMs在公式规范定义阶段通常表现良好,但在数据检索和数值执行阶段的准确率大幅下降;DR智能体的表现始终优于配备搜索功能的LLMs,不过在实际金融分析场景中仍存在不可靠性。这些发现为开发更强大、更可信的金融领域DR智能体提供了参考。
英文摘要
Financial indicators are essential tools for transforming raw financial data into interpretable measures for various downstream tasks, such as valuation, risk assessment, and economic analysis. However, existing financial benchmarks largely focus on answer-level accuracy and often assume that relevant data are already provided, leaving the assessment of the intermediate process of indicator construction underexplored. In this work, we propose FinDeepIndicator, the first benchmark dedicated to evaluating Deep Research (DR) agents in end-to-end financial indicator construction. Specifically, FinDeepIndicator evaluates DR agents across four stages in indicator construction: formula specification, data collection, indicator calculation, and answer generation, and covers fundamental, technical, and macroeconomic indicators organized into 21 fine-grained sub-categories. It contains 3,350 curated question-answer (QA) pairs derived from both U.S. and Chinese markets, 10 years of historical financial data, and 800 listed companies. Extensive experiments on search-equipped Large Language Models (LLMs) and DR agents show that, while LLMs generally perform well in formula specification, their accuracy drops substantially during data retrieval and numerical execution. DR agents consistently outperform search-equipped LLMs, yet remain unreliable in realistic financial analysis settings. These findings provide insights for developing more capable and trustworthy DR agents in finance.