AI 中文总结
针对时间序列基础模型强化学习后训练的次优崩溃问题,提出真实值邻域正则化方法,可缓解该问题并提升模型性能,且能灵活集成到各类强化学习方法中。
AI 中文摘要
时间序列预测(TSF)在众多实际应用中发挥着重要作用。近期,在大规模数据集上预训练的时间序列基础模型(TSFMs)展现出强大的泛化能力,成为TSF领域的重要范式。强化学习(RL)后训练作为进一步提升其下游任务性能的手段,受到越来越多关注。然而,研究发现,在某些预测区域,RL后训练可能会逐渐使TSFMs的输出分布偏离真实值,从而限制其性能,该现象被称为“次优崩溃”。分析表明,初始难以在真实值附近采样高质量轨迹是导致次优崩溃的重要因素。为解决这一问题,本文提出面向TSFMs RL后训练的真实值邻域正则化(GTN-R)方法,该方法以真实值为参考定位高质量区域,引导模型的概率质量向真实值邻域移动,提高高质量轨迹的采样概率,缓解次优崩溃并提升性能。此外,GTN-R可灵活集成到TSFMs的各类RL方法中,大量实验验证了其有效性。
英文摘要
Time series forecasting (TSF) plays an important role in a wide range of real-world applications. Recently, time series foundation models (TSFMs), pretrained on large-scale datasets, have demonstrated strong generalization capabilities and emerged as an important paradigm for TSF. Reinforcement learning (RL) post-training has consequently attracted growing attention as a means of further improving their performance on downstream tasks. However, we find that, in certain forecast regions, RL post-training may gradually shift the output distributions of TSFMs away from the ground truth, thereby limiting their performance. We refer to this phenomenon as \textbf{suboptimal collapse}. Our analysis suggests that difficulty in initially sampling high-quality trajectories near the ground truth is an important contributing factor to suboptimal collapse. To address this issue, we propose Ground-Truth Neighborhood Regularization (GTN-R) for RL post-training of TSFMs. GTN-R uses the ground truth as a reference for locating high-quality regions and guides the model's probability mass toward the ground-truth neighborhood. This increases the probability of sampling high-quality trajectories, mitigates suboptimal collapse, and improves performance. Moreover, GTN-R can be flexibly integrated into various RL methods for TSFMs. Extensive experiments show its effectiveness.