arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LEAP:面向基于大语言模型的概率预测的似然 elicitation 与聚合

LEAP: Likelihood Elicitation and Aggregation for LLM-based Probabilistic Forecasting

Yufei Chen, Yiran Zhao, Xiaogang Xu, Qipeng Xie, Jiafei Wu, Zhe Liu

arXiv 2609.01337首次发表:更新:

发表机构

Shandong University; Nanjing University of Aeronautics and Astronautics; Zhejiang University; Ningbo Global Innovation Center, Zhejiang University; The Hong Kong University of Science and Technology (Guangzhou)(山东大学; 南京航空航天大学; 浙江大学; 浙江大学宁波科创中心; 香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对基于LLM的整体预测设计模糊证据影响、合并不确定性的问题,提出LEAP方法,构建多任务基准并验证其可提升预测与校准指标,性能更稳定。

AI 中文摘要

基于大语言模型(LLM)的预测系统已在金融市场、体育结果等现实任务中取得性能提升,这主要得益于更强大的搜索能力与工具使用能力。许多此类系统仍要求 LLM 读取所有收集到的证据后共同生成最终预测,我们将这种设计称为“整体预测(Monolithic Prediction)”。该设计会模糊单个证据项对结果的影响,还会合并不同竞争结果间的不确定性。我们提出 LEAP(Likelihood Elicitation and Aggregation for Probabilistic forecasting,即概率预测的似然 elicitation 与聚合),它重新组织了预测阶段对收集到的证据的使用方式。LEAP 会分别检查每个证据项,引出描述其对目标影响的似然参数;随后,显式先验与确定性概率模型会将这些似然组合为后验分布。该流程支持连续型、单选择、多选择预测,同时保留可复现的证据贡献。我们构建了涵盖预测、信息检索、浏览任务的基准,并在自研智能体循环及多个智能体 CLI 框架上评估 LEAP。在相同证据条件下,LEAP 在各类模型上提升了多数预测与校准指标,且在对先验访问、推理预算、聚合方式的控制对比中表现更优。

英文摘要

LLM-based forecasting systems have improved on real-world tasks such as financial markets and sports outcomes, largely through stronger search and tool use. Many systems still ask an LLM to read all collected evidence together and produce the final forecast. We call this design Monolithic Prediction. It can obscure how individual evidence items affect the result and collapse uncertainty across competing outcomes. We propose LEAP (Likelihood Elicitation and Aggregation for Probabilistic forecasting), which reorganizes how collected evidence is used in the prediction stage. LEAP examines each evidence item separately and elicits likelihood parameters that describe its implications for the target. An explicit prior and a deterministic probabilistic model then combine these likelihoods into a posterior distribution. This procedure supports continuous, single-choice, and multi-choice forecasts while preserving reproducible evidence contributions. We build a benchmark covering forecasting, information-seeking, and browsing tasks, and evaluate LEAP on our own agent loop and several agent CLI frameworks. Given the same evidence, LEAP improves most prediction and calibration metrics across models and remains stronger under controlled comparisons of prior access, inference budget, and aggregation.

CommentsAccepted to the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026). 20 pages, 3 figures, and 15 tables. Code: https://github.com/layingfish/LEAP

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑