发表机构
Fudan University; Microsoft Research; Nankai University; Jilin University(复旦大学; 微软研究院; 南开大学; 吉林大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对单一时间序列预测模型的局限性,提出基于LLM推理的REATS集成方法,通过三方面设计实现样本自适应可解释加权,在八个基准测试上优于竞争基线,具备强泛化与迁移能力。
AI 中文摘要
由于现实世界时间序列的多样性,没有单一预测模型能在所有样本上始终占据优势。集成学习通过结合互补的模型优势解决这一问题,但现有方法依赖固定规则或仅基于数值输入的黑箱模型,未能利用大语言模型(LLM)的推理能力做出可解释的加权决策。我们提出REATS,它将LLM的推理能力作为智能集成路由,共同处理文本化的时间模式描述和数值特征,通过思维链推理生成可解释的、样本自适应的集成权重。为实现基于LLM的有效集成,我们研究了其关键设计选择并提出:(i)结构化输入流水线,将原始时间序列转换为具有固定token成本的混合文本-数值表示,无需依赖API即可构建基于规则的思维链,并补充检索到的相似样本先验;(ii)多样化的多行权重监督方案,结合token高效的百分比表格格式,降低数值复杂度并缓解LLM的幻觉问题;(iii)结合监督微调(SFT)与生成式偏好优化(GRPO)的两阶段微调框架,其中倒数奖励映射将连续无界的均方误差(MSE)差距转换为有界信号,放大接近最优模型的灵敏度,解决基于回归的GRPO中固有的均匀灵敏度和离群值主导的优势压缩问题。在八个基准测试上的实验表明,REATS优于竞争性集成基线,同时提供自然语言解释,并展现出对未见过的候选模型的强迁移学习和域外泛化能力。
英文摘要
Due to the diversity of real-world time series, no single forecasting model consistently dominates across all samples. Ensemble learning addresses this by combining complementary model strengths, yet existing methods rely on fixed rules or black-box models based solely on numerical inputs, failing to leverage LLM reasoning for interpretable weighting decisions. We propose REATS, which leverages LLM reasoning capabilities as an intelligent ensemble router that jointly processes textual temporal pattern descriptions and numerical features to produce interpretable, sample-adaptive ensemble weights through chain-of-thought reasoning. To enable effective LLM-based ensembling, we study its key design choices and propose: (i) a structured input pipeline that transforms raw time series into hybrid textual--numerical representations with fixed token cost, enabling rule-based chain-of-thought construction without API dependency, augmented with retrieved similar-sample priors; (ii) a diverse multi-row weight supervision scheme coupled with a token-efficient percentage-table format that reduces numerical complexity and mitigates LLM hallucinations; and (iii) a two-stage fine-tuning framework combining SFT with GRPO, where a reciprocal reward mapping transforms the continuous unbounded MSE gap into bounded signals with amplified near-oracle sensitivity, addressing the uniform sensitivity and outlier-dominated advantage compression inherent in naive reward designs for regression-based GRPO. Experiments on eight benchmarks demonstrate that REATS outperforms competitive ensemble baselines while providing natural language explanations and demonstrating strong transfer learning and out-of-domain generalization to unseen candidate models.