发表机构
University of Zurich; ETH Zurich; University of Cambridge(苏黎世大学; 苏黎世联邦理工学院; 剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出参数高效框架,将预训练语言模型适配于时间序列预测,通过连续块嵌入优于文本提示,仅更新不到1%参数即达专业架构精度。
AI 中文摘要
我们通过一个参数高效的迁移学习框架,研究将预训练语言模型适配到单变量时间序列预测的问题,旨在理解哪些设计选择能驱动有效的跨模态迁移。尽管语言模型作用于离散的文本词元,时间序列却由具有时间依赖性的连续数值观测组成。为弥合这一模态差距,我们将固定长度的时间序列块直接投影到预训练GPT-2骨干网络的嵌入空间中,绕开文本词元化,将Transformer视为通用序列编码器。通过在涵盖能源、天气、交通和金融的七个基准数据集上进行受控消融研究,我们分析了以下因素的影响:(i)表示策略(连续嵌入对比文本序列化);(ii)适配机制(冻结骨干网络对比部分或完全微调);(iii)架构组件,如适配器、池化策略和预测头;(iv)输入上下文长度。基于连续块的嵌入始终优于文本提示和随机初始化的骨干网络。适配后的流程在更新不到总模型参数1%的情况下,达到了专业预测架构的MASE范围。结果进一步表明,冻结预训练骨干网络并训练轻量级投影和适配器模块,提供了良好的精度-效率权衡,且在不同上下文长度下行为稳定。
英文摘要
We study the adaptation of pretrained language models to univariate time-series forecasting through a parameter-efficient transfer learning framework, with the goal of understanding which design choices drive effective cross-modal transfer. While language models operate on discrete textual tokens, time series consist of continuous numerical observations with temporal dependencies. To bridge this modality gap, we project fixed-length time-series patches directly into the embedding space of a pretrained GPT-2 backbone, bypassing textual tokenization and treating the Transformer as a generic sequence encoder. Through controlled ablation studies on seven benchmark datasets spanning energy, weather, traffic, and finance, we analyze the effects of (i)~representation strategy (continuous embeddings versus textual serialisation), (ii)~adaptation regime (frozen backbone versus partial or full fine-tuning), (iii)~architectural components such as adapters, pooling strategies, and prediction heads, and (iv)~input context length. Continuous patch-based embeddings consistently outperform textual prompting and randomly initialised backbones. The adapted pipeline attains MASE within the range of specialised forecasting architectures while updating less than 1\% of total model parameters. Results further indicate that freezing the pretrained backbone and training lightweight projection and adapter modules provides a favourable accuracy--efficiency trade-off with stable behaviour across varying context lengths.