发表机构
Johns Hopkins University(约翰斯·霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究构建金融情感分类基准,对比TF-IDF朴素贝叶斯等模型,发现QLoRA可提升Qwen2.5性能,但分类准确率高的模型在收益预测的经济有效性上无显著优势,揭示分类准确率与可交易信号的差距。
AI 中文摘要
金融情感分类器通常基于人工标签进行评估,但强大的语言性能并不一定意味着具有经济实用的收益可预测性。本研究通过两项实验将这两个问题分开。首先,我们从5个金融文本数据集构建了一个统一的三分类基准,比较了TF-IDF朴素贝叶斯、现成的FinBERT和Financial-RoBERTa编码器、零样本Qwen2.5-7B,以及采用QLoRA适配的Qwen2.5-7B、LLaMA3-8B和Mistral-7B模型。Mistral-7B取得了最佳测试准确率(0.8840)和宏F1值(0.8771),而QLoRA将Qwen2.5的宏F1值从0.7274提升至0.8615。逆频率类加权损失并未改善Qwen2.5的性能。其次,我们在时间上独立的2019年Benzinga样本上评估经济有效性,该样本包含10637条独特标题和13115条标题-股票观测值,对应固定的标普100(S&P 100)标的范围。模型概率被转换为连续情感得分,按股票和信号日期聚合,并与下一交易日的1天、2天、3天和5天收益对齐。所有7个下游模型在1天 horizon 产生正但较小的平均秩信息系数,最大的为FinBERT的0.0143。在Newey-West推断和错误发现率校正后,28项模型-horizon测试均不显著。投资组合结果也未能为表现最佳的分类器确立稳健优势。研究结果表明,QLoRA对金融情感适配是有效的,同时也揭示了分类准确率与可交易的横截面信号之间存在明显差距。
英文摘要
Financial sentiment classifiers are commonly evaluated against human labels, but strong linguistic performance does not necessarily imply economically useful return predictability. This study separates these questions through two experiments. First, we construct a unified three-class benchmark from five financial text datasets and compare TF--IDF Naive Bayes, off-the-shelf FinBERT and Financial-RoBERTa encoders, zero-shot Qwen2.5-7B, and QLoRA-adapted Qwen2.5-7B, LLaMA3-8B, and Mistral-7B models. Mistral-7B achieves the best test accuracy (0.8840) and macro-F1 (0.8771), while QLoRA raises Qwen2.5's macro-F1 from 0.7274 to 0.8615. An inverse-frequency class-weighted loss does not improve Qwen2.5. Second, we evaluate economic validity on a temporally separate 2019 Benzinga sample containing 10,637 unique headlines and 13,115 headline--stock observations for a fixed S\&P~100 universe. Model probabilities are converted into continuous sentiment scores, aggregated by stock and signal date, and aligned with next-session returns over one-, two-, three-, and five-day horizons. All seven downstream models produce positive but small mean rank information coefficients at the one-day horizon; the largest is 0.0143 for FinBERT. None of the 28 model--horizon tests remains significant after Newey--West inference and false-discovery-rate correction. Portfolio results likewise fail to establish a robust advantage for the best-performing classifiers. The findings show that QLoRA is effective for financial sentiment adaptation, while also documenting a clear gap between classification accuracy and tradable cross-sectional signals.
CommentsWaiting to submit to a conference (ICAIF)