arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17223cs.CLcs.LG

金融新闻NLP中的时间泄露:针对特定 regime 并购信号的多架构审计

Temporal Leakage in Financial News NLP: A Multi-Architecture Audit with a Regime-Specific M&A Signal

Chenhao Xue, Raslen Guesmi, Siwei Feng, Yucheng Gong, Jacob Xavier Sundram, Jordan Pang, Lan Wang, Julian Kaljuvee

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对金融新闻NLP的时间泄露问题,在多模型架构上审计了其影响,发现并购信号具有时间局限性,主张将泄露审计作为金融NLP基准的必备披露项。

中文摘要 AI 辅助

金融新闻方向预测已成为热门的NLP基准任务,但报告的性能提升高度依赖训练-测试划分是时间顺序划分还是随机划分,即时间泄露问题。我们在包含49799篇文章的语料库上,针对TF-IDF、MiniLM、FinBERT、微调后的RoBERTa-large、DeBERTa-v3-large,以及Llama-3和Qwen2.5大模型的零/少样本、LoRA探针,共16种特征-模型组合审计了这种依赖关系:随机划分使马修斯相关系数(MCC)提升了1.1倍至6.5倍,且与模型容量和特征丰富度正相关,端到端FinBERT微调会重新放大而非缩小该差距(规模匹配下的比值为1.75倍)。按事件类型划分,并购(M&A)是唯一在近乎时间顺序评估下具有锁定测试正信号的审计类别(TF-IDF的MCC:仅训练集为0.138,训练集+验证集重新拟合后为0.068;10000次置换检验p<10⁻³);该信号无法迁移到FNSPID的2009-2020年美国语料库,说明该标题信号局限于我们2024-2025年偏向欧洲的并购语义,而非通用预测因子。三名独立角色标注员一致认为收购方标注的文章是信号来源,这是受限于能力的定性收敛,而非经假设检验的不对称性。时间顺序划分在金融NLP中的作用,类似于资产定价中的特征清洗:它剥离了新闻中可预测的陈旧成分,留下的残差很小、事件局部化且词汇层面较浅。我们主张将泄露审计作为金融NLP基准的必备披露项。

英文摘要

Financial-news direction prediction has become a popular NLP benchmark, yet reported gains depend critically on whether the train-test split is chronological or random, i.e., on temporal leakage. We audit this dependence on a 49,799-article corpus across 16 feature-model combinations spanning TF-IDF, MiniLM, FinBERT, and fine-tuned RoBERTa-large / DeBERTa-v3-large, plus separate zero/few-shot and LoRA probes of Llama-3 and Qwen2.5 LLMs: random splits inflate MCC by $1.1\times$ to $6.5\times$, tracking model capacity and feature richness, and end-to-end FinBERT fine-tuning re-amplifies rather than closes the gap (size-matched ratio $1.75\times$). Conditioning on event type, mergers and acquisitions (M&A) is the only audited category with a positive locked-test signal under near-temporal chronological evaluation (TF-IDF MCC $= 0.138$ train-only, $0.068$ under train$\cup$val refit; 10,000-permutation $p < 10^{-3}$); the signal does not transfer to FNSPID's 2009-2020 U.S. corpus, localising the headline to our 2024-2025 European-tilted M&A semantics rather than a universal predictor. Three independent role labellers converge on acquirer-tagged articles as the signal locus, a power-limited qualitative convergence rather than a hypothesis-tested asymmetry. Chronological splitting plays for financial NLP the role characteristics-purging plays for asset pricing: it strips the predictable, stale component of news and leaves a residual that is small, event-localized, and lexically shallow. We advocate leakage audits as a required disclosure for financial-NLP benchmarks.

发表机构

  • Predictive Labs Ltd(Predictive实验室有限公司)
  • University of Oxford(牛津大学)
  • Imperial College London(伦敦帝国学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑