MM-FinEval:面向真实世界金融预测的多任务多模态基准
MM-FinEval: A Multi-Task Multimodal Benchmark for Real-World Financial Forecasting
浏览论文内容
中文总结 AI 辅助
本文提出MM-FinEval,一个包含2019-2022年2045场标普500财报电话会议、覆盖12个金融任务的多模态基准,验证三模态输入对多模态大模型金融预测的有效性。
中文摘要 AI 辅助
从财报电话会议中进行金融预测,要求模型能够对复杂的企业披露信息、市场预期以及微妙的沟通信号进行推理。然而,现有的金融基准通常局限于单模态输入或单任务设置,这使得评估多模态大语言模型(LLMs)能否支持真实世界的金融分析变得困难。在本文中,我们引入了MM-FinEval,这是一个新颖的基准,旨在跨多个金融任务评估多模态大语言模型。MM-FinEval涵盖了2019年至2022年的多样化时间线。整个提出的数据集包含2,045场标普500指数(S&P 500)成分股公司的财报电话会议作为输入,以及12个金融任务标签作为输出。每个输入包含三种模态:电话会议的逐字文本记录、会议期间使用的相应演示幻灯片,以及完整的音频录音。为了建立严格的评估框架,我们分析了19个基线模型,涵盖三种不同的模型类别:图像-文本、音频-文本以及任意到任意(Any-to-Any)配置。我们观察到,处理所有三种模态的小型任意到任意模型取得了强劲的性能,即使与受限于双模态输入的更大规模专有模型相比也是如此。这表明我们的三模态数据集设计引入了有用且非冗余的信息。这些结果验证了文本、音频和视觉数据作为重要的互补信号,模拟了人类专家分析师在决策过程中的表现。
英文摘要
Financial forecasting from earnings conference calls requires models to reason over complex corporate disclosures, market expectations, and subtle communication signals. However, existing financial benchmarks are often limited to unimodal inputs or single-task settings, making it difficult to evaluate whether multimodal large language models (LLMs) can support real-world financial analysis. In this paper, we introduce MM-FinEval, a novel benchmark designed to evaluate multimodal LLMs across multiple financial tasks. MM-FinEval spans a diverse timeline from 2019 to 2022. The entire proposed dataset contains 2,045 S\&P 500 conference earning calls as inputs and 12 financial task labels as outputs. Each input contains three modalities: a word-to-word text transcript of the earning call, the corresponding presentation slides used during the call, and the entire audio recording. To establish a rigorous evaluation framework, we analyze 19 baseline models across three distinct model categories: Image-Text, Audio-Text, and Any-to-Any configurations. We observe that small-size Any-to-Any models processing all three modalities achieve strong performance, even when compared against larger proprietary models restricted to two-modality inputs. This indicates that our tri-modal dataset design introduces useful, non-redundant information. These results validate that text, audio, and visual data serve as important, complementary signals that mimic the decision-making process of expert human analysts.
发表机构
- Northwestern University(西北大学)
- New Jersey Institute of Technology(新泽西理工学院)
- Georgia Institute of Technology(佐治亚理工学院)
- Pace University(佩斯大学)
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。