相同文本,不同数字:基于LLM的度量指标的分歧
Same Text, Different Numbers: The Divergence of LLM-Based Measures
- EDHEC Business School(EDHEC商学院)
- University of Groningen(格罗宁根大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究检验了七个LLM在十三种文本度量上的跨模型一致性,发现秩相关仅0.52,模型选择显著影响下游推断,故LLM变量应视为模型依赖的测量并需跨提供商验证。
AI中文摘要:
研究者越来越多地使用生成式大语言模型(LLM)将公司文本转换为实证变量。我们使用十三种度量指标,包括情感、管理层清晰度、不确定性、回答具体性以及气候和政治风险,考察了基于LLM的文本度量在多大程度上对模型选择具有不变性。来自不同提供商的七个LLM对标准普尔500指数公司的财报电话会议记录进行了这些构念的评分。跨模型秩相关平均仅为0.52,而各提供商共有的转录本层面差异仅占总得分变异的34%。跨模型分歧并不能预测随后的分析师或市场分歧,这与存在大量模型特定成分而非底层披露中的共同模糊性相一致。模型选择显著影响下游推断,系数大小、符号和统计显著性在各模型间差异很大。对大多数构念,跨提供商平均使转录本排名更稳定,但得分水平仍对集成中包含的模型敏感。因此,LLM生成的变量应被视为依赖于模型的测量结果,并需跨提供商进行验证。
英文摘要:
Researchers increasingly use generative large language models (LLMs) to convert corporate text into empirical variables. We examine the extent to which LLM-based textual measures are invariant to model choice using thirteen measures, including sentiment, management clarity, uncertainty, answer specificity, and climate and political risk. Seven LLMs from different providers score earnings call transcripts of S&P 500 companies on these constructs. Cross-model rank correlations average only 0.52, and transcript-level differences common across providers account for only 34% of total score variation. Cross-model disagreement does not predict subsequent analyst or market disagreement, consistent with a substantial model-specific component rather than common ambiguity in the underlying disclosure. Model choice significantly affects downstream inference, with coefficient magnitudes, signs, and statistical significance varying substantially across models. Averaging across providers makes transcript rankings more stable for most constructs, but score levels remain sensitive to the models included in the ensemble. LLM-generated variables should therefore be treated as model-contingent measurements and validated across providers.