arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19378cs.LGcs.CVstat.ML

通过输入依赖的长卷积实现原生多维次二次算子

Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions

David R. Wessels, Farhad Ramezanghorbani, Alireza Moradzadeh, David W. Romero, Olivia Viessmann, Maksim Zhdanov, John St. John, Ken Janik, David M Knigge, Yuche… 展开作者

David R. Wessels, Farhad Ramezanghorbani, Alireza Moradzadeh, David W. Romero, Olivia Viessmann, Maksim Zhdanov, John St. John, Ken Janik, David M Knigge, Yucheng Tang, Erik J Bekkers, Saee Gopal Paliwal

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对多维数据应用注意力机制次二次替代方案的权衡问题,提出HyenaND算子,通过特定卷积直接作用于多维数据原生几何结构,CUDA实现nSubQ加速,实验表明其在多领域表现出色,纯堆栈匹配基线,混合配置更优。

中文摘要 AI 辅助

注意力机制的次二次替代方案在应用于多维数据时需要权衡:标准卷积缺乏全局感受野和输入依赖性,而循环模型需要将图像、体数据和偏微分方程(PDE)等数据光栅化为违反其空间结构的临时一维扫描顺序。我们引入了HyenaND,这是一种次二次、全局、输入依赖的算子,它通过与隐式参数化的全局、输入依赖的多维卷积核进行卷积,直接作用于多维数据的原生几何结构。我们的CUDA实现nSubQ融合了FFT卷积路径,将HyenaND的O(L log L)缩放转换为实际加速。在长上下文基因组学、计算机视觉、医学成像和PDE建模中,纯HyenaND堆栈与强大的注意力基线的准确性相匹配,而交错HyenaND和注意力层的混合配置优于纯注意力和基于强循环的混合配置。

英文摘要

Subquadratic alternatives to attention require compromises when applied to multi-dimensional data: standard convolutions lack global receptive fields and input dependency, while recurrent models require rasterizing data such as images, volumes, and partial differential equation (PDE) into an ad-hoc $1\rm D$ scan order that violates their spatial structure. We introduce \textit{HyenaND}, a subquadratic, global, input-dependent operator that acts directly on the native geometry of multidimensional data through convolutions with implicitly parametrized global, input-dependent multi-dimensional convolutional kernels. Our CUDA implementation, \texttt{nSubQ}, fuses the FFT-convolution path to turn HyenaND's $\mathcal{O}(L \log L)$ scaling into wall-clock speedups. Across long-context genomics, computer vision, medical imaging, and PDE modeling, pure HyenaND stacks match the accuracy of strong attention baselines, while hybrid configurations that interleave HyenaND and attention layers outperform both pure attention and strong recurrence-based hybrids.

发表机构

  • AMLab, University of Amsterdam(阿姆斯特丹大学AMLab)
  • NVIDIA(英伟达)
  • Cartesia AI(Cartesia人工智能公司)
  • New Theory AI(新理论人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

↑