通过上下文感知的NLP流水线将专家审议转化为金融信号
Converting Expert Deliberation into Financial Signals Through A Context-Aware NLP Pipeline
- Franklin Templeton Investments(富兰克林邓普顿投资公司)
- Blend360(Blend360公司)
- Santa Clara University(圣克拉拉大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出CDSP流水线将投资委员会会议记录转化为金融特征,经实验其结合句子嵌入与CDSP特征的模型预测准确率达73%,证实专家审议含前瞻性金融信息。
AI中文摘要:
我们提出CDSP(上下文条件审议信号流水线),将投资委员会的会议记录转化为结构化预测特征。CDSP将会议记录分割为主题块,利用大语言模型(LLM)分配资产类别上下文标签,将金融关键词映射到预定义的标签分类体系,并构建互补特征:情感极性和提及频率。该特征工程框架应用于包含48次月度委员会会议的数据集,以预测下月全球股票是否将表现优于全球债券。在使用工程特征、原始会议记录文本、句子嵌入及组合表示的实验中,预测准确率介于62%至73%之间,而始终选择股票的策略表现优于债券的概率为60.4%。表现最佳的模型(准确率73%)结合了句子嵌入与工程CDSP特征,取得0.73的F1分数(尽管与始终选择股票的策略相比无统计学显著性)。对于多个分类体系类别,情感比提及频率携带更强的信号。这些发现表明,专家的审议可能包含可通过上下文感知NLP提取的前瞻性信息。
英文摘要:
We introduce the CDSP (context-conditional deliberation signal pipeline), converting an investment committee's meeting transcripts into structured predictive features. CDSP segments the meeting transcripts into topical chunks, assigns asset-class context labels using a large language model (LLM), maps financial keywords to a pre-determined taxonomy of labels, and constructs complementary features: sentiment polarity and mention frequency. This feature engineering framework is applied to a dataset spanning 48 monthly committee meetings to predict if global equities will perform better or worse than global bonds in the following month. In experiments with engineered features, raw transcript text, sentence embeddings, and combined representations, the prediction accuracy ranges from 62% to 73%, compared to always choosing stocks, which outperforms bonds 60.4% of the time. The best (73% accurate) model combines sentence embeddings with engineered CDSP features, achieving a 0.73 F1 score (although this is not statistically significant compared to always choosing stocks). Sentiment carries a stronger signal than mention frequency for several taxonomy categories. These findings suggest that experts' deliberations may contain forward-looking information that context-aware NLP can extract.