从词频到语境:资产定价中的主题模型
From Word Counts to Context: Topic Models for Asset Pricing
- University of California, Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本研究比较LDA与句子转换器在新闻文本中构建系统性风险因素的效果,发现后者在主题连贯性和财务表现上更优,组合模型夏普比率达1.03。
中文摘要 AI 辅助
新闻可能揭示系统性风险,但其语境是否有助于增强系统性风险因素的构建仍不清楚。我们试图测试,在从非结构化文本数据生成的主题词列表的连贯性方面,使用句子转换器是否比潜在狄利克雷分配(LDA)等技术有所改进。为测试这一点,我们对同一组包含394,661篇文章的非结构化文本数据以及相同的下游金融投资组合构建流程进行了应用,仅文本层有所不同,包括每个模型使用的文章文本长度以及主题词的排序方式:我们将LDA与使用k均值聚类的冻结句子转换器进行基准比较。我们发现,句子转换器分支在连贯性(以NPMI衡量)和财务表现(以夏普比率衡量)方面均观察到更高的得分,尽管现有测试并未确立其优越性。进一步的探索性设定,如利用球形聚类和多期风险暴露,组合模型的超额收益夏普比率为1.03。我们认为,在非结构化新闻文本上应用上下文感知技术具有一定的前景,但可能需要使用仅基于每个日期可用信息的更严格测试以及更广泛的数据集来增强对观察到的表现的信心。
英文摘要
News may reveal systematic risk, but whether its context enhances the construction of systematic risk factors is still unclear. We seek to test whether utilizing a sentence transformer represents an improvement over techniques such as Latent Dirichlet Allocation (LDA) in the coherence of topic term lists generated from unstructured text data. To test this, the same collection of unstructured text data comprising of 394,661 articles and the same downstream financial portfolio construction pipeline were applied with the text layer differing, including the length of article text each model used and how topic terms were ranked: we benchmark LDA against a frozen sentence transformer with k-means clustering. We find that the sentence transformer branch had higher observed scores both in terms of coherence (measured by NPMI) as well as financial performance (measured by Sharpe), although the available tests do not establish outperformance. Further exploratory specifications such as utilizing spherical clustering and multi-horizon exposures had an observed excess-return Sharpe of 1.03 for the combined model. We believe that there is some promise in applying context-aware techniques on unstructured news text, but stricter tests using only information available at each date and broader datasets may be required to enhance the confidence in the observed performance.