arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

序数门限、基数投注:将大语言模型(LLM)置信度匹配至金融决策算子

Ordinal Gates, Cardinal Bets: Matching LLM Confidence to the Financial Decision Operator

Rayansh Singh, Sara Rezaeimanesh

arXiv 2609.00187首次发表:更新:

发表机构

Michigan State University(密歇根州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对金融决策场景,发现LLM置信度的映射与尺度需联合评估,匹配二者可显著提升确定等价收益,且效果不受单一模型驱动。

AI 中文摘要

大语言模型(LLM)的置信度分数并非可独立部署的对象:其决策价值取决于使用它们的下游算子和暴露控制器。单调 recalibration(重新校准)无法改变覆盖率匹配的基于排名的门限,而头寸规模(position sizing)会利用分数幅度,因此改变置信度映射可能会使适配于先前分数分布的尺度失效。我们针对纳斯达克100指数股票的FactSet新闻开展测试,在2021年数据上拟合映射和尺度,于2022-2023年对9个开放权重LLM进行样本外评估。将原始映射和正确性映射与独立拟合的尺度交叉应用,结果显示两个组件单独不可迁移:尺度迁移使8/9个模型的确定等价收益(CER)降低,并产生大量风险目标误差。将每个映射与其拟合尺度匹配,在冻结尺度控制下使集成模型的CER每年提升9.2个百分点(p<0.001),且在排除贡献最大的单个模型后,效果仍显著(每年+5.5个百分点),因此并非由单一案例驱动。然而,在相同的自适应波动率控制器下,增量效果降至每年+1.6个百分点,存在显著的控制器交互作用。每年的滚动窗口效应更小,尽管映射-尺度交互在每个折叠中均为正。因此,置信度变换应与使用它们的下游控制器联合评估。

英文摘要

LLM confidence scores are not independently deployable objects: their decision value depends on the downstream operator and exposure controller that consume them. Monotone recalibration cannot change a coverage-matched rank-based gate, whereas position sizing consumes score magnitude, so changing a confidence map can invalidate a scale fitted to the previous score distribution. We test this on FactSet news for Nasdaq-100 equities, fitting maps and scales on 2021 and evaluating nine open-weight LLMs out-of-sample on 2022--2023. Cross-applying raw and correctness maps with independently fitted scales shows that the two components are not portable alone: scale transfer reduces certainty-equivalent return (CER) in $8/9$ models and produces large risk-target errors. Matching each map with its fitted scale improves ensemble CER by $9.2$ percentage points per year under frozen-scale control ($p<0.001$), and the effect remains significant when the single largest-contributing model is excluded ($+5.5$pp/yr), so it is not driven by one case. Under an identical adaptive-volatility controller, however, the incremental effect falls to $+1.6$pp/yr, with a significant controller interaction. Annual walk-forward effects are smaller, although map--scale interaction remains positive in every fold. Confidence transformations should therefore be evaluated jointly with the downstream controllers that consume them.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑