发表机构
Institut Universitaire de France(法国高等研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究探讨高斯混合模型上Softmax注意力的无限提示极限,表明其可通过梯度方法学习统计任务最优解,兼具线性注意力的线性任务处理能力与依赖查询的非线性任务解决能力。
AI 中文摘要
Softmax注意力是Transformer的核心,已展现出卓越能力,但其底层机制仍仅被部分理解。近期理论工作研究高斯提示,其中无限提示极限将Softmax注意力简化为线性映射,但也消除了使其区别于线性注意力的依赖查询的选择机制。本研究探讨高斯混合模型上Softmax注意力的无限提示极限,该模型既保留高斯数据的易处理性,又引入潜在结构、多模态性与非线性依赖。我们表明,Softmax注意力可通过基于梯度的方法表示并学习一系列统计任务的最优解,包括监督分类与去噪。研究结果凸显Softmax注意力的两种互补能力:它能像更简单的线性注意力对应物一样有效恢复线性任务,同时还能利用依赖查询的上下文选择能力解决线性注意力无法处理的非线性任务。
英文摘要
Softmax attention, at the heart of Transformers, has demonstrated remarkable capabilities. Yet its underlying mechanisms remain only partially understood. Recent theoretical work studies Gaussian prompts, where the infinite-prompt limit reduces softmax attention to a linear map, but also removes the query-dependent selection that distinguishes it from linear attention. This work studies the infinite-prompt limit of softmax attention on Gaussian mixtures, which retain the tractability of Gaussian data while introducing latent structure, multimodality, and nonlinear dependencies. We show that softmax attention can represent and learn, via gradient-based methods, optimal solutions to a range of statistical tasks, including supervised classification and denoising. Our results highlight two complementary capabilities of softmax attention: it can recover linear tasks as effectively as its simpler linear counterpart, while also exploiting query-dependent context selection to solve nonlinear tasks beyond the reach of linear attention.