arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14174cs.LGq-fin.CPq-fin.MFq-fin.ST

10-K报告的哪些部分重要?全文与风险因素情绪的聚合相关价值

How Much of a 10-K Matters? Aggregation-Dependent Value of Full-Text versus Risk-Factor Sentiment

Sanggyu Sean Choi

首次发表
浏览论文内容

中文总结 AI 辅助

研究探讨10-K文件中全文与风险因素情绪的聚合相关价值,将监督词典学习方法扩展到10-K文件及其风险因素部分,在多级别训练情绪分数,评估指标表现,发现不同聚合级别表现差异及词典基线问题,确立后续系统的情绪生成方法。

中文摘要 AI 辅助

金融情绪提取主要依赖新闻文本和仅针对回报标签的监督提取,而10-K文件以及波动性(目标风险披露最适合提供信息的对象)相对未被充分探索。我们将一种监督词典学习方法扩展到10-K文件及其第1A项风险因素部分,在行业、投资组合和单个公司三个聚合级别针对回报和波动性标签训练情绪分数。在来自94家纳斯达克100指数科技成分股(2006 - 2023年)的1383份文件中,我们评估了由此产生的12个情绪指标在分类准确性、与实际市场结果的相关性以及定性词汇内容方面的表现。全文本在行业和投资组合层面针对两个目标产生更准确的情绪,但在单个公司层面则相反,第1A项较窄的部分表现更好,我们将这种效应归因于文档量与每个聚合级别可用的独立训练信号量之间的相互作用。Loughran - McDonald词典基线在每个测试级别与价格始终呈强负相关,强调了监督方法对监管披露文本的价值。这些发现以及它们推动的设计选择,确立了后续更大规模多源系统的情绪生成方法。

英文摘要

Financial sentiment extraction has largely relied on news text and supervised extraction against return labels alone, leaving 10-K filings -- and volatility, the target risk disclosure is arguably best suited to informing -- comparatively unexplored. We extend a supervised lexicon-learning approach to 10-K filings and their Item 1A risk-factor sections, training sentiment scores against both return and volatility labels at three levels of aggregation: sector, portfolio, and individual firm. Across 1,383 filings from 94 Nasdaq-100 technology constituents (2006--2023), we evaluate the resulting twelve sentiment metrics on classification accuracy, correlation with realised market outcomes, and qualitative lexical content. Full-filing text produces more accurate sentiment at the sector and portfolio level for both targets, but this reverses at the individual-firm level, where the narrower Item 1A section performs better -- an effect we attribute to the interaction between document volume and the amount of independent training signal available at each level of aggregation. A Loughran-McDonald dictionary baseline is consistently, strongly negatively correlated with price at every level tested, underscoring the value of a supervised approach for regulatory disclosure text. These findings, and the design choices they motivate, establish the sentiment-generation methodology underlying a subsequent, larger-scale, multi-source system.

↑