arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

渐进式记忆Transformer:面向时间序列的记忆感知注意力

Progressive Memory Transformer: Memory-Aware Attention for Time-Series

Tord Sture Stangeland, Andreas Köhler, Steffen Mæland, Adín Ramíres Rivera

arXiv 2609.31351首次发表:更新:

发表机构

University of Oslo; NORSAR; University of Tromsø; Western Norway University of Applied Sciences(奥斯陆大学; 挪威地震观测与研究机构; 特罗姆瑟大学; 挪威西部应用科学大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出渐进式记忆Transformer(PMT),通过窗口对齐的可写记忆在局部、中程和全局三个尺度上显式建模时间序列结构层次,在七个分类基准和预测任务上实现强低标签分类与竞争性预测性能。

AI 中文摘要

时间序列同时在多个尺度上携带结构(细粒度变化、中程基序和全局属性),下游任务也在相应不同的尺度上运行。大多数现有的自监督学习方法通过实例级对比损失和有限的时序邻域监督来全局监督表示,但并未显式利用结构层次。我们提出一种学习框架,显式地在三个尺度上独立强制结构层次:用于标记连续性的局部目标、用于窗口级基序的中程目标,以及用于序列级一致性的全局目标。实现该框架需要主干网络在每个尺度上暴露表示;我们引入渐进式记忆Transformer(PMT),它用可写入的、窗口对齐的记忆增强Transformer,该记忆在传统Transformer已提供的标记和序列级表示之外,还暴露中程尺度。在七个UCR/UEA/UCI分类基准、一个线索保留探针和预测基准上,PMT学习到的表示在全局、中程和局部尺度上探测良好——强低标签分类(1–5%标签)、跨多个预测视野的竞争性预测性能,以及记忆状态捕获中程基序的定量和定性证据。

英文摘要

Time-series carry structure simultaneously at multiple scales (fine-grained variation, mid-range motifs, and global properties) and downstream tasks operate at correspondingly different scales. Most existing self-supervised learning approaches supervise representations globally via instance-level contrastive losses and limited temporal neighborhood supervision, but do not explicitly exploit the structural hierarchy. We propose a learning framework that explicitly enforces a structural hierarchy across three scales independently: a local objective for token continuity, a mid-range objective for window-level motifs, and a global objective for sequence-level agreement. Realizing this framework requires the backbone to expose a representation at each scale; we introduce \textbf{Progressive Memory Transformer} (PMT), which augments a transformer with writable, window-aligned memory that exposes the mid-range scale alongside the token and sequence-level representations conventional transformers already provide. Across seven UCR/UEA/UCI classification benchmarks, a cue-retention probe, and forecasting benchmarks, PMT learns representations that probe well at the global, mid-range, and local scales---strong low-label classification (1--5\% labels), competitive forecasting performance across multiple horizons, and quantitative and qualitative evidence that memory states capture mid-range motifs.

CommentsTo appear in NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑