arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

同日同故事,提前一天不同信号:金融情绪的双重有效性

Human Agreement and Return Association Are Not Interchangeable Criteria

AS Aravinthakshan, Laven Srivastava, Harsh Nandwani

arXiv 2609.11144首次发表:更新:

发表机构

Manipal Institute of Technology; Perssonify(马尼帕尔理工学院; Perssonify)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过证券集体诉讼语料库检验金融情绪工具的双重有效性,发现基准一致性确立语义效度但不决定预测排名,消息量无法预测市场影响。

AI 中文摘要

金融自然语言处理有一个标准工作流程:先根据人工标注验证情绪工具,然后信任其提取市场信号。这假设了两种评估衡量的是同一事物。我们在一个可以同时测量两者的环境中检验这一假设:一个证券集体诉讼语料库(2002-2025年),将70,500条X消息与异常股票回报相关联,并附有单一标注者的人工标注黄金样本。通过五个工具(VADER、Loughran-McDonald、FinBERT、Twitter-RoBERTa和一个LLM标注器)在一条相同的流水线上运行,我们发现构念效度与预测效度之间的关系取决于采样惯例和分数表示。在常规的特定方法采样下,人工一致性更接近分级同日关联,而非提前一天的领先关系。然而,在固定样本量面板上,一致性在两个时间跨度上具有相似的分级秩相关,而粗略排序仍然较弱。因此,基准一致性确立了语义效度,但本身并不决定预测排名。在一个垃圾信息占17.6%的对话中,消息量既不能预测市场损害,也不能预测和解金额。

英文摘要

Financial NLP has a standard workflow: validate a sentiment tool against human labels, then trust it to extract market signal. This assumes the two evaluations measure the same thing. We test that assumption in a setting where both can be measured at once: a corpus of securities class actions (2002-2025) linking 70,500 X messages to abnormal stock returns, with a single-annotator human labelled gold sample. Running five instruments (VADER, Loughran-McDonald, FinBERT, Twitter-RoBERTa, and an LLM annotator) through one identical pipeline, we find that the relationship between construct and predictive validity depends on the sampling convention and score representation. Under conventional method-specific sampling, human agreement aligns more closely with graded same-day associations than with one-day leads. On a fixed-n panel, however, agreement has similar graded rank correlations at both horizons, while the coarse ordering remains weak. Benchmark agreement therefore establishes semantic validity but does not by itself determine predictive rankings. In a conversation that is 17.6% spam, message volume predicts neither market damage nor settlement size.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑