arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20304q-fin.RM

LLM金融预测中校准诱导的退化:关于次日市场风险的审计跟踪案例研究

Calibration-Induced Degeneracy in LLM Financial Forecasting: An Audit-Trailed Case Study on Next-Day Market Risk

Arin Mohanty

AI总结:

该研究以次日市场风险为对象,发现LLM金融预测中存在校准诱导的退化问题,提出校准可行性检查点,验证低成本基准的关键作用并终止了无效的全历史推理阶段。

AI中文摘要:

只有当校准允许大型语言模型(LLM)特征影响预测时,这些成本高昂的特征才具有实际意义。我们在两项大盘基金的次日风险研究中发现了这一关联的失效。研究采用全历史评分先于2022年校准,校准后四个LLM权重均被设为零,因此后续的856个评分无法对评估产生影响,我们将此现象称为校准诱导的退化。允许带符号权重后,所有四个映射关系均被重新激活,但经家庭式错误校正后,无一能提升预测效果。相比之下,一项近乎零成本的标题计数将SPY的方差预测损失降低了0.001720(95%家庭式区间:[0.000719, 0.002830]),因此该低成本基准是关键诊断工具。我们提出校准可行性检查点:拟合映射、在预设校准值上扰动特征、并要求在获取留存特征前存在有意义的预测响应,该检查无需使用留存结果,在此场景下本可终止付费的全历史推理阶段。

英文摘要:

Costly LLM features matter only if calibration lets them affect the forecast. We document a failure of this link in a next-day risk study of two broad-market funds. Full-history scoring preceded the 2022 calibration. Calibration then set all four LLM weights to zero. The 856 later scores therefore could not affect the evaluation. We call this calibration-induced degeneracy. Allowing signed weights reactivated all four mappings. None improved forecasts after familywise correction. By contrast, a near-zero-cost headline count reduced SPY variance-forecast loss by 0.001720 (95 percent familywise interval: [0.000719, 0.002830]). The cheap baseline is therefore a critical diagnostic. We propose a calibration-viability checkpoint. Fit the mapping, perturb the feature over prespecified calibration values, and require a meaningful forecast response before acquiring holdout features. The check uses no holdout outcomes. Here, it would have stopped the paid full-history inference phase.

补充信息

↑