arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于大语言模型(LLM)的预测与决策智能体:方法、训练、评估及应用

LLM-based Agents for Forecasting and Prediction: Methods, Training, Evaluation, and Applications

Xiaogang Xu, Jiaqi Tang, Jianmin Chen, Yingying Yan, Zhenchao Tang, Xiangxin Zhou, Xiaobin Hu, Wei Wei, Jinfeng Wu, Qifeng Chen, Lu Zhou, Jiafei Wu, Zhe Liu, Jianwei Yin, Weimin Zheng

arXiv 2608.23058首次发表:更新:

AI 中文总结

本文综述基于LLM的预测智能体的架构、训练与评估方法,分析其局限并介绍多领域应用,指出未来需解决分布偏移校准等问题。

AI 中文摘要

当前大语言模型(LLM)可支撑将语言推理与时间数据、证据检索、外部工具及迭代预测相结合的预测系统。本文研究基于LLM的预测智能体,即语言模型对未来或当前未观测目标生成带评分的预测结果的系统。我们将架构分为三类:独立LLM工作流对编码后的时间序列或事件上下文进行处理;工具与检索增强型智能体整合外部证据;混合系统将LLM与统计模型或基础模型配对。随后我们综述训练方法与评估协议,同时考察正反两方面证据,包括对微小输入扰动的敏感性、LLM组件未提升准确率的 ablation 实验、可能反映数据污染而非时间推理的基准增益。我们涵盖其在金融、气象、健康、能源及运营领域的应用,并总结评估所用的基准与数据集。现有证据表明,测量是核心局限,未来研究需解决分布偏移下的校准、抗污染的在线评估、成本与准确率的联合明确报告,以及处理部署预测与被预测结果间反馈的方法。

英文摘要

Large language models (LLMs) now support forecasting systems that combine language-based reasoning with temporal data, evidence retrieval, external tools, and iterative prediction. We investigate LLM-based forecasting agents, meaning systems in which a language model contributes to a scored prediction about a future or currently unobserved target. We organize architectures into three groups. Standalone LLM workflows operate on encoded time series or event context. Tool- and retrieval-augmented agents incorporate external evidence. Hybrid systems pair LLMs with statistical or foundation models. We then review training methods and evaluation protocols. We examine negative as well as positive evidence, including sensitivity to small input perturbations, ablations in which the LLM component does not improve accuracy, and benchmark gains that may reflect contamination instead of temporal reasoning. We cover applications in finance, weather, health, energy, and operations, and we summarize the benchmarks and datasets used for evaluation. The evidence indicates that measurement is a central limitation. Future work requires calibration under distribution shift, contamination-resistant live evaluation, explicit reporting of cost and accuracy together, and methods for handling feedback between deployed forecasts and the outcomes being forecast.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑