arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Tucker 瓶颈注意力用于多维序列建模

Tucker Bottleneck Attention for Multi-Dimensional Sequence Modeling

Ryan Solgi, Parsa Madinei, Zheng Zhang

arXiv 2610.09090首次发表:更新:

发表机构

University of California-Santa Barbara(加州大学圣塔芭芭拉分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出 Tucker 瓶颈注意力(TuBA),利用低秩张量结构实现次二次计算的高效全局 token 混合,在视频预测和天气预报中显著降低误差与计算成本。

AI 中文摘要

自注意力的二次成本限制了其扩展到来自多维数据的长序列的可扩展性。我们引入了 Tucker 瓶颈注意力(TuBA),它利用低秩张量结构实现高效的全局 token 混合。TuBA 将隐藏张量投影为紧凑的 Tucker 核,在核上执行多头自注意力和线性投影,并将更新写回环境空间,从而实现次二次计算。其自回归扩展结合了核内的双向交互与跨核的因果注意力。在视频预测和全球天气预报中,TuBA 相比标准和高效注意力以及任务特定模型实现了有利的精度-效率权衡。与标准自注意力相比,TuBA 在视频预测中减少了高达 24.7% 的误差和 66.6% 的计算量,在自回归天气预报中减少了 37.1% 的误差和 85.1% 的计算量,加速高达 4.27 倍。低秩 Tucker 核和多帧生成也分别优于全秩注意力和逐帧生成。

英文摘要

The quadratic cost of self-attention limits scalability to long sequences from multidimensional data. We introduce Tucker bottleneck attention (TuBA), which exploits low-rank tensor structure for efficient global token mixing. TuBA projects hidden tensors into compact Tucker cores, performs multi-head self-attention and linear projections on the cores, and writes updates back to the ambient space, enabling subquadratic computation. Its autoregressive extension combines bidirectional interactions within cores with causal attention across cores. On video prediction and global weather forecasting, TuBA achieves favorable accuracy-efficiency trade-offs over standard and efficient attention and task-specific models. Compared to standard self-attention, TuBA reduces error and computation by up to 24.7% and 66.6% for video prediction and 37.1% and 85.1% for autoregressive weather forecasting, with speedups up to 4.27 times. Low-rank Tucker cores and multi-frame generation also outperform full-rank attention and frame-by-frame generation, respectively.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑