发表机构
IU International University of Applied Sciences(IU国际应用科学大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
比较基于词典和基于大语言模型的情绪分析在表情包股票尾部风险检测中的应用,构建情绪指标并评估与市场回报关系,结果显示基于LLM的指标表征更丰富,但与市场走势关系因资产而异。
AI 中文摘要
本文对基于词典和基于大语言模型(LLM)的情绪分析进行实证比较,以从高波动股票市场的社交媒体话语中提取与市场相关的信号。利用来自r/WallStreetBets的Reddit数据并聚焦表情包股票,构建时间对齐的情绪指标并评估其与市场回报的关系。基于LLM的方法生成多维情绪表征,而基线依赖基于VADER词典的模型。通过多种分析方法评估两种方法,结果表明基于LLM的指标提供了更丰富的多维表征,但与市场走势的关系在不同资产间存在异质性。
英文摘要
This paper presents an empirical comparison of lexicon-based and Large Language Model (LLM)-based sentiment analysis for extracting market-relevant signals from social media discourse in highly volatile equity markets. Using Reddit data from r/WallStreetBets and focusing on meme stocks (GME, AMC, NOK), we construct time-aligned sentiment indicators and evaluate their relationship with market returns, with particular attention to extreme positive return events in the upper tail of the return distribution. The LLM-based approach generates multidimensional sentiment representations capturing emotional polarity, bullishness, sarcasm likelihood, and topical relevance, whereas the baseline relies on the VADER lexicon-based model. We evaluate both approaches using lead/lag correlation analysis, OLS regression, ROC-AUC-based directional classification, and a quantile-based early-warning framework. The results indicate that LLM-derived indicators provide a richer multidimensional representation and exhibit stronger asset-specific statistical structure than the lexicon-based baseline. However, their relationship with market movements remains heterogeneous across assets, suggesting that increased linguistic expressiveness does not necessarily translate into stable forecasting performance in retail-driven volatility regimes.
Comments10 pages, 1 figure, 4 tables