AI 中文总结
研究针对大语言模型在金融预测中证据不足时高置信度的问题,提出FinAbstain框架,通过多模态检索增强生成及选择性预测,经多种不确定性评估方法和指标评估,贡献了时间安全架构、复合不确定性公式和可重复评估蓝图。
AI 中文摘要
大语言模型在合成金融叙事时,证据不足、过时或矛盾时可能会表现出高置信度,这在预测中后果严重。我们提出了FinAbstain,一个用于不确定性校准多模态检索增强生成(RAG)的研究框架,具有选择性预测功能。一个时间点检索器仅允许预测时间戳公开的信息,并为基本面、新闻、技术、风险和验证代理提供特定模态的证据。它们的概率评估通过检索相关性、证据矛盾性、重复样本一致性和历史校准统计进行汇总。在一个通用的时间顺序协议下评估了温度缩放、等距回归、共形预测和一种提出的混合不确定性分数。一个控制器仅在不确定性低于验证阈值时预测看涨、看跌或中性结果;否则弃权、请求证据、减少风险敞口或将案例提交人工审查。评估涵盖了一日和五日异常收益方向、二十日波动率区间和弃权决策,使用了准确性、校准、风险覆盖、引用、交易、延迟和成本指标。为了在完整数据收集完成之前使设计可审计,我们报告了明确标记的模拟结果。这些结果说明了预期的假设:校准弃权可能以覆盖范围换取更低的选择性误差和回撤。贡献在于一个时间安全的架构、一个复合不确定性公式和一个用于基于证据的选择性金融预测的可重复评估蓝图。
英文摘要
Large language models (LLMs) can synthesize financial narratives but may express high confidence when evidence is sparse, stale, or contradictory. This failure is especially consequential in forecasting, where filings, news, prices, volume, and technical signals can disagree. We present FinAbstain, a research framework for uncertainty-calibrated multimodal retrieval-augmented generation (RAG) with selective prediction. A point-in-time retriever admits only information public at the forecast timestamp and supplies modality-specific evidence to fundamental, news, technical, risk, and verification agents. Their probabilistic assessments are aggregated with retrieval relevance, evidence contradiction, repeated-sample consistency, and historical calibration statistics. Temperature scaling, isotonic regression, conformal prediction, and a proposed hybrid uncertainty score are evaluated under a common chronological protocol. A controller predicts bullish, bearish, or neutral outcomes only when uncertainty is below a validated threshold; otherwise it abstains, requests evidence, reduces exposure, or routes the case to human review. The evaluation covers one- and five-day abnormal-return direction, twenty-day volatility intervals, and abstention decisions, using accuracy, calibration, risk--coverage, citation, trading, latency, and cost metrics. To make the design auditable before a full data collection is complete, we report explicitly labeled simulated results rather than empirical claims. These results illustrate the intended hypothesis: calibrated abstention may trade coverage for lower selective error and drawdown. The contribution is a time-safe architecture, a composite uncertainty formulation, and a reproducible evaluation blueprint for evidence-grounded selective financial forecasting.