发表机构
School of Automation Science and Electrical Engineering, Beihang University; Hangzhou International Innovation Institute, Beihang University; State Key Laboratory of Intelligent Manufacturing System Technology; School of Computer Science and Artificial Intelligence, Zhengzhou University(北京航空航天大学自动化科学与电气工程学院; 北京航空航天大学杭州国际创新研究院; 智能制造系统技术国家重点实验室; 郑州大学计算机科学与人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对工业时间序列单模态建模局限及连接时间序列与文本语义的挑战,提出VLT多模态基础模型,通过设计Time-MoE等机制联合建模多种模态,经实验验证其在多种复杂设置下优于现有方法,提升了鲁棒性和泛化能力。
AI 中文摘要
工业时间序列是预测与健康管理(PHM)的基础,可确保航空发动机等工业设备的可靠性和安全性。但现有方法多限于单模态建模,限制了其在复杂场景中的泛化。尽管大语言模型的进展为多模态学习带来新机遇,但连接连续时间序列信号和离散文本语义仍是挑战。为此提出VLT,一个联合建模时间序列、频谱视觉表示和文本知识的多模态基础模型。关键在于利用频谱作为视觉桥梁连接连续时间信号和离散语义。具体设计了时间感知专家混合模型(Time-MoE)捕捉异构时间动态,频率-文本增强学习者在共享表示空间中联合建模频谱和语义特征。还引入以时间为中心的梯度对齐机制减轻跨模态优化冲突。在多个工业数据集上的大量实验表明,VLT优于现有方法,在少样本、有噪声和不完全模态设置下具有卓越的鲁棒性和泛化能力。
英文摘要
Industrial time series serve as the foundation for Prognostics and Health Management (PHM) to ensure the reliability and safety of industrial equipment such as aero-engines. However, existing approaches are typically limited to single-modality modeling, which restricts their generalization in complex scenarios. Although recent advances in large language models (LLMs) provide new opportunities for multimodal learning, bridging continuous time-series signals and discrete textual semantics remains an open challenge. To this end, we propose VLT, a multimodal foundation model that jointly models time-series, frequency-spectrum visual representations, and textual knowledge. A key insight is to utilize the frequency spectrum as a visual bridge to connect continuous temporal signals with discrete semantics. Specifically, a Time-aware Mixture-of-Experts (Time-MoE) is designed to capture heterogeneous temporal dynamics, while a Frequency-Text Augmented Learner enables joint modeling of spectral and semantic features within a shared representation space. Furthermore, a time-centric gradient alignment mechanism is introduced to mitigate cross-modal optimization conflicts via gradient normalization and reliability-aware dynamic reweighting. Extensive experiments on multiple industrial datasets demonstrate that VLT outperforms state-of-the-art methods, achieving superior robustness and generalization under few-shot, noisy, and incomplete-modality settings.
Comments18 pages, 13 figures, and 13 tables, including supplementary material. Haiteng Wang and Jingheng Yan contributed equally to this work