发表机构
Loyola University of Chicago; North Dakota State University(芝加哥洛约拉大学; 北达科他州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文评估十二种情绪模型,发现通用大语言模型分类性能与金融专用模型相当,但高准确率未转化为更强的经济关系,情绪度量与盈利意外相关却无法解释短期市场反应。
AI 中文摘要
金融情绪度量在实证金融中被广泛使用,但通用大语言模型(LLM)是否优于现有的金融专用方法仍不明确。本文评估了十二种情绪模型,包括基于词典的方法、金融专用变换器模型和开源大语言模型,使用两个标准:语言有效性和经济有效性。我们发现,通用大语言模型在无需任务特定微调的情况下,实现了与金融专用变换器模型相当的分类性能。然而,更高的分类准确率并未转化为更强的经济关系。几种模型产生的情绪度量与盈利意外显著相关,但没有任何模型与次日股票回报显著相关。模型性能在盈利大幅超出或低于预期的公告中表现最强,而在盈利意外较为温和的公告中表现明显较弱。这些发现表明,金融情绪捕捉了企业经济表现的信息,但解释短期市场反应的能力有限。
英文摘要
Financial sentiment measures are widely used in empirical finance, but it remains unclear whether general-purpose large language models (LLMs) improve on existing finance-specific methods. This paper evaluates twelve sentiment models, including dictionary-based methods, finance-specific transformers, and open-source LLMs, using two criteria: linguistic validity and economic validity. We find that general-purpose LLMs achieve classification performance comparable to finance-specific transformer models without task-specific fine-tuning. However, higher classification accuracy does not translate into stronger economic relationships. Several models produce sentiment measures that are significantly associated with earnings surprises, but none is significantly associated with next-day stock returns. Model performance is strongest for announcements with large earnings beats or misses and substantially weaker for announcements with more moderate earnings surprises. These findings suggest that financial sentiment captures information about firms' economic performance but has limited ability to explain short-run market reactions
Comments19 pages, 5 figures, 3 tables, preprint