arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自进化时间序列预测智能体:基于情景记忆与在线策略学习

Self-Evolving Time-Series Forecasting Agents with Episodic Memory and Online Policy Learning

Junyi Wang, Yilin Wang, Wen Wu, Chao Zhang

arXiv 2609.32689首次发表:更新:

AI 中文总结

针对现有预测智能体忽视历史反馈的问题,提出FASE,结合情景记忆与在线策略学习,利用已完成实例反馈自我进化,在29个配置上MAE降低9.1%。

AI 中文摘要

基于大语言模型的智能体越来越多地被用于时间序列预测,因为它们能够组织上下文信息、执行多步分析,并引导完成预测任务所需的行动序列。现有的大多数智能体仅关注当前的预测实例。然而,在实际部署中,预测通常以在线方式运行,随着预测起点向前推进,新的预测基于当前可用的历史数据发出,而早期实例的真实目标值也逐步变得可用。这些目标值为早期实例中采取的行动提供了反馈,但现有智能体通常不会保留或利用这些信息来调整其后续行动。为解决这一局限,我们引入了FASE,一种反馈感知的自进化预测智能体,它将此类反馈转化为针对后续预测实例的任务特定经验。FASE结合了情景记忆(用于检索相关的已完成实例)与在线策略学习(将跨实例累积的反馈总结为排序指导)。所提出的框架在从GIFT-Eval基准中选取的29个数据集配置上进行了评估。在这29个配置中,FASE在评估方法中取得了最强的总体点预测性能,并将归一化平均绝对误差(MAE)相对于最佳个体基础模型基线降低了9.1%。结果进一步表明,随着延迟反馈的累积,FASE的累积优势逐渐增加。这些发现共同表明,FASE能够通过已完成预测实例的反馈持续自我进化,而无需更新大语言模型的参数。

英文摘要

LLM-based agents are increasingly used for time-series forecasting because they can organise contextual information, perform multi-step analysis, and guide the sequence of actions required to complete forecasting tasks. Most existing agents focus only on the current forecasting instance. However, in real-world deployments, forecasting commonly operates online, with new forecasts issued from the currently available history as the forecast origin advances and the ground-truth targets of earlier instances progressively become available. These targets provide feedback on the actions taken in earlier instances, yet existing agents generally do not preserve or utilise this information to adapt their subsequent actions. To address this limitation, we introduce FASE, a Feedback-Aware Self-Evolving forecasting agent that converts such feedback into task-specific experience for subsequent forecasting instances. FASE combines episodic memory, which retrieves relevant completed instances, with online policy learning, which summarises the feedback accumulated across instances into ranking guidance. The proposed framework is evaluated on 29 dataset configurations selected from the GIFT-Eval benchmark. Across these 29 configurations, FASE attains the strongest aggregate point forecasting performance among the evaluated methods and reduces the normalised MAE by 9.1% relative to the best individual foundation model baseline. The results further indicate that the cumulative advantage of FASE increases as delayed feedback accumulates. Together, these findings demonstrate that FASE can continually self-evolve through feedback from completed forecasting instances without updating the parameters of the LLM.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑