DisclosureBeta:一种基于LLM读取的风险披露的制度条件贝塔的测量通道理论
DisclosureBeta: A Measurement-Channel Theory for Regime-Conditioned Betas from LLM-Read Risk Disclosures
AI总结:
本文提出DisclosureBeta理论,将LLM建模为风险特征测量通道,构建制度条件贝塔估计方法,自适应结合文本与滚动窗口估计量,填补现有方法的理论与误差预算空白。
AI中文摘要:
当一家公司的价格历史过短而无法信赖时(例如S-1备案公司、近期上市主体或刚经历制度断裂的主体),交易台所需的贝塔估计是一个亟待解决的问题。现有最优方法退化为可比公司的贝塔,且无误差预算;近期基于文本的同类方法Breitung(2025)虽在IPO样本中展现出强经验准确性,但缺乏识别理论、误差预算及下界。本文填补了这一空白:将大型语言模型(LLM)建模为对公司潜在风险特征的噪声测量通道,并将该通道噪声纳入资产定价的误差预算。在分段平稳的Fama-French五因子模型中,因子载荷是潜在风险特征与推断制度的函数。我们在关于通道、检测器及制度内抽样的明确假设下,证明了制度条件载荷函数的可识别性与一致性,并给出匹配下界,表明对于仅观测到收益、因子、LLM特征及制度估计的任何估计量,披露噪声与检测器误分类项是不可避免的。一项披露激励推论表明,估计精度与公司层面披露激励度量(DIM)呈单调关系。基于文本与滚动窗口估计量的自适应凸组合始终不劣于任一组成部分,且在价格历史过短、陈旧或跨越已检测制度断裂时,会将权重转向文本。对价格历史薄弱公司的冻结预注册面板的经验评估即将开展;本预印本记录了理论与预注册设计,以独立于经验结果确立优先权。
英文摘要:
The problem is the beta a desk needs when a firm's price history is too short to trust: an S-1 filer, a recent listing, or a name just past a regime break. The state of the art collapses to a comparable-firm peer beta with no error budget, and the recent text-based competitor Breitung (2025) reports strong empirical IPO accuracy but no identification theory, no error budget, and no lower bound. We fill that gap. We model a large language model as a noisy measurement channel on a firm's latent risk characteristics and write its channel noise into the asset-pricing error budget. In a piecewise-stationary Fama-French five-factor model the loadings are a function of latent risk characteristics and an inferred regime. We prove identification and consistency of the regime-conditional loading function under explicit assumptions on the channel, the detector, and within-regime sampling, and give a matching lower bound showing that the disclosure-noise and detector-misclassification terms are unavoidable for any estimator that observes only returns, factors, LLM features, and a regime estimate. A disclosure-incentive corollary makes estimation precision monotone in a firm-level disclosure-incentive measure (DIM). An adaptive convex combination of the text-based and rolling-window estimators is never worse than either component and shifts its weight toward text exactly when price history is short, stale, or straddles a detected regime break. The empirical evaluation on a frozen, pre-registered panel of price-history-thin firms is forthcoming; this preprint records the theory and the pre-registered design so priority is established independently of the empirical outcome.