arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30695cs.LG

液体门控注意力

Liquid Gated Attention

  • College of Computer Science and Technology, Jilin University(吉林大学计算机科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

Yiheng Jiang, Yuanbo Xu, Yongjian Yang

中文总结 AI 辅助

针对现实时间序列的不规则采样与长范围建模需求,提出无求解器的并行时间算子LGA,构建线性复杂度的LFormer,在多任务多数据集上实现长程依赖建模等能力,性能具竞争力且效率线性。

中文摘要 AI 辅助

现实世界的时间序列常表现出不规则采样和扩展的时间范围,要求模型能在任意区间内捕捉连续时间动态,且不会产生过高的扩展成本。离散时间方法将可变时间区间压缩为固定位置步长;依赖求解器的连续时间模型保留时间结构,但依赖顺序积分,无法并行化;而无求解器的近似方法避免了这一成本,但 none(此处保留英文)未将观测时间区间与输入驱动的状态调制相结合。我们提出液体门控注意力(Liquid Gated Attention, LGA),一种无求解器的并行时间算子。通过用观测时间区间参数化输入驱动的门控机制,LGA引入了连续时间归纳偏置,并将隐藏状态演化公式化为快速权重关联记忆,实现了时间维度上的并行计算。利用非因果编码中的矩阵结合性和因果编码中的前缀扫描,LGA在两种模式下均达到了与序列长度成线性关系的时间复杂度。序列级归一化对累积时间衰减进行约束,以实现稳定的长范围优化。基于LGA,我们实例化了LFormer,一个用于连续时间表示学习的模块化骨干网络。在覆盖多达17984步的6项任务和16个数据集上,LFormer展现出长程依赖建模、细粒度状态跟踪以及从稀疏和噪声观测中重建轨迹的能力,同时在与最先进的离散时间和连续时间基线相比时,达到了具有线性扩展效率的竞争力性能。

英文摘要

Real-world time series often exhibit irregular sampling and extended temporal horizons, requiring models to capture continuous-time dynamics across arbitrary intervals without prohibitive scaling costs. Discrete-time methods collapse variable time intervals into static positional steps; solver-dependent continuous-time models preserve temporal structure but rely on sequential integration, precluding parallelization; and solver-free approximations avoid this cost yet none couples observed time intervals with input-driven state modulation. We propose Liquid Gated Attention (LGA), a solver-free parallel temporal operator. By parameterizing an input-driven gating mechanism with observed time intervals, LGA introduces a continuous-time inductive bias and formulates hidden state evolution as a fast-weight associative memory, enabling parallel computation across the temporal dimension. Using matrix associativity in non-causal encoding and a prefix scan in causal encoding, LGA attains linear temporal complexity in sequence length in both modes. A sequence-level normalization bounds cumulative temporal decay for stable long-horizon optimization. Building on LGA, we instantiate LFormer, a modular backbone for continuous-time representation learning. Across six tasks and sixteen datasets spanning up to 17,984 steps, LFormer demonstrates long-range dependency modeling, fine-grained state tracking, and trajectory reconstruction from sparse and noisy observations, while delivering competitive performance against state-of-the-art discrete-time and continuous-time baselines with linear scaling efficiency.

补充信息

↑