发表机构
UT Austin(得克萨斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
介绍用于金融预测的FinBench基准,它通过时间门控避免前瞻偏差,用适当评分规则评估。任务要求模型输出正回报概率和预测区间,用多种分数评估,通过小型试点运行展示校准敏感指标能区分不同预测行为。
AI 中文摘要
大语言模型越来越多地被用作观察、规划和行动的智能系统组件。在金融领域,一旦其输出用于确定交易规模或分配风险,即使是“辅助”系统也会与决策相关。一个关键的失败模式是置信度与能力差距:一个仅略优于随机猜测但持续过度自信的模型,在典型的投注规模规则下,将产生负的长期增长。现有基准强调语义理解或点准确性,但未在定义真实市场的时间约束和非平稳性下直接测试概率校准。我们引入了FinBench,这是一个旨在评估金融预测校准和不确定性质量的基准,它(i)严格设置时间门控以避免前瞻性偏差,(ii)使用严格适当的评分规则进行评估,对虚假置信度进行惩罚。FinBench任务要求模型输出(a)正回报的概率和(b)实现对数回报的80%预测区间;评估使用Brier分数、Winkler区间分数以及针对硬基线的技能分数。本文描述了基准规范,并报告了一次小型试点运行(一个交易日;三个流动性股票代码;33次预测)作为对流程的合理性检查。试点说明了校准敏感指标如何区分“自信但脆弱”的行为和不确定性感知预测。
英文摘要
Large language models (LLMs) are increasingly used as components of agentic systems that observe, plan, and act. In finance, even "assistive" systems become decision-relevant once their outputs are used to size trades or allocate risk. A key failure mode is the confidence--competence gap: a model that is only slightly better than chance but consistently overconfident will, under typical bet-sizing rules, generate negative long-run growth. Existing benchmarks emphasize semantic understanding or point accuracy, but do not directly test probabilistic calibration under the temporal constraints and non-stationarity that define real markets. We introduce FinBench, a benchmark designed to evaluate calibration and uncertainty quality for financial forecasting in a setting that is (i) strictly time-gated to avoid look-ahead bias and (ii) evaluated with strictly proper scoring rules that penalize hallucinated confidence. FinBench tasks require models to output (a) a probability of positive return and (b) an 80% prediction interval for realized log return; evaluation uses the Brier score and the Winkler interval score, along with skill scores against hard baselines. This paper describes the benchmark specification and reports a small pilot run (one trading day; three liquid tickers; 33 forecasts) as a sanity check of the pipeline. The pilot illustrates how calibration-sensitive metrics distinguish between "confident but fragile" behavior and uncertainty-aware forecasting.
Comments9 pages, 3 figures