arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新思考基于门控线性注意力和KAN的视觉架构

Rethinking Vision Architectures with Gated Linear Attention and KAN

Ali Mehizel, Oussama Khaldi

arXiv 2609.22506首次发表:更新:

发表机构

ENSSEA; Université Paris Cité(ENSSEA; 巴黎西岱大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出LKAT架构,结合门控线性注意力与KAN前馈网络,在ImageNet-100和CIFAR上超越ViT等基线,验证了二者作为互补归纳偏置的有效性。

AI 中文摘要

视觉Transformer将大部分参数分配给多层感知机(MLP)用于通道混合,而token交互通常依赖于二次复杂度的多头自注意力(MHSA)。线性注意力将序列复杂度降低到O(N),但仍与softmax Transformer中相同的固定激活函数MLP耦合。Kolmogorov-Arnold网络(KAN)则将可学习的单变量映射置于边上,但先前的视觉KAN要么保留MHSA,要么完全省略注意力。我们引入LKAT(线性Kolmogorov-Arnold Transformer),一种各向同性的ViT编码器,将分块门控线性注意力(GLA)与两层KAN前馈网络耦合,并为径向基网格映射提供了I/O感知的融合RBF-KAN内核。在共享的DeiT风格训练方案下,我们将LKAT与ViT、ViT-5和MLP-Mixer进行比较。LKAT-B在ImageNet-100上超过了ViT-B/16、ViT-5-B和Mixer-B/16。Tiny/Small/Base规模的LKAT变体在CIFAR-10/100上表现出一致的扩展性,且ImageNet-100预训练可迁移到CIFAR微调。实验结果支持门控线性注意力和基于KAN的径向基函数作为中等规模视觉表示学习的互补归纳偏置。

英文摘要

Vision Transformers devote most of their parameters to MLPs for channel mixing, but still rely on quadratic multi-head self-attention for token interactions. While linear attention fixes the complexity problem, bringing it down to O(N), it is usually just paired with the same fixed-activation MLP as before. Kolmogorov-Arnold Networks take a different approach, placing learnable univariate functions on the edges instead. However, existing vision KANs either retain standard attention or remove attention entirely, so the two ideas have not been effectively combined. We introduce LKAT (Linear Kolmogorov-Arnold Transformer) to close this gap: an isotropic ViT-style encoder that couples chunk-wise Gated Linear Attention with a two-layer KAN feed-forward block, backed by an I/O-aware fused RBF-KAN kernel to make radial-basis grid functions efficient in practice. Under a shared DeiT-style training recipe, LKAT-B outperforms ViT-B/16, ViT-5-B, and Mixer-B/16 on ImageNet-100, while Tiny, Small, and Base variants scale consistently on CIFAR-10/100. ImageNet-100 pretraining also transfers effectively to CIFAR fine-tuning, suggesting that gated linear attention and KAN-based radial basis functions provide complementary inductive biases for mid-scale visual representation learning. Code: https://github.com/mehizelali/linear-kan-transformer

Comments19 pages, 9 figures. Code available at https://github.com/mehizelali/linear-kan-transformer

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑