arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SAGE:面向视觉-语言时间序列预测的逐变量语义增强方法

SAGE: Variate-Wise Semantic Augmentation for Vision-Language Time Series Forecasting

Haizhao Fan, Xinyi Le

arXiv 2608.26829首次发表:更新:

发表机构

Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对时间序列预测模型缺乏语义知识的问题,提出基于CLIP的SAGE框架,通过双重利用CLIP实现多模态监督,在8个长期基准及M4数据集上取得最优精度。

AI 中文摘要

时间序列预测模型仅基于原始数值序列进行运算,缺乏领域专家隐含利用的语义知识,例如各变量的物理含义、统计特性及时间动态。近期弥合该差距的研究分为两类:部分方法在推理阶段依赖大语言模型(LLM),计算成本高昂;其他方法在数据集层面应用统一文本提示,忽略了各变量间的异质语义。本文提出SAGE(基于接地编码的感知与增强),这是一种基于CLIP的端到端框架,可联合建模时间、跨变量、文本及视觉信息。CLIP文本编码器处理频率增强的片段和变量标记,门控残差路径注入变量特定描述与统计描述符;并行地,冻结的CLIP视觉编码器通过仅训练阶段的对比目标,将渲染序列与时间表示对齐。这种CLIP的双重使用在预测循环中不引入LLM,同时补充了语义与视觉监督。在8个长期基准及M4数据集上,SAGE达到了最优的预测精度; ablation实验证实了多模态对齐与变量级知识带来的互补增益。

英文摘要

Time series forecasting models operate on raw numerical sequences, lacking the semantic knowledge that domain experts implicitly leverage, such as the physical meaning of each variable, its statistical behavior, and its temporal dynamics. Recent efforts to bridge this gap fall into two camps. Some rely on large language models at inference time, which is computationally expensive. Others apply uniform textual prompts at the dataset level, ignoring the heterogeneous semantics across individual variates. We propose SAGE (Seeing and Augmenting with Grounded Encoding), an end-to-end CLIP-based framework that jointly models temporal, cross-variable, textual, and visual information. The CLIP text encoder processes frequency-enhanced patches and variable tokens, while gated residual paths inject variable-specific descriptions and statistical descriptors. In parallel, the frozen CLIP vision encoder aligns rendered series with temporal representations through a training-only contrastive objective. This dual use of CLIP adds complementary semantic and visual supervision without placing an LLM in the forecasting loop. Across eight long-term benchmarks and M4, SAGE achieves state-of-the-art accuracy. Ablations confirm complementary gains from multimodal alignment and variable-level knowledge.

Comments10 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑