arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

线性化2-单纯形注意力机制

Linearized 2-Simplicial Attention

Aritra Das, Dhruman Gupta, Debayan Gupta

arXiv 2608.09307首次发表:更新:

发表机构

Truth Audit Labs(Truth Audit Labs)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出线性化2-单纯形注意力机制,结合Kimi Delta Attention构建无softmax注意力的模型,在匹配计算下获最高平均下游准确率,16k上下文时提升准确率并降低LAMBADA困惑度。

AI 中文摘要

我们提出了一种线性化形式的2-单纯形注意力机制,方法是将三线性得分重写为复合查询与键之间的内积,使得在一个令牌轴上的求和与普通softmax注意力的形式相同。随后,我们用正随机特征近似该求和,并将全部历史存储在固定大小的状态中,而第二个轴在近期令牌的短窗口内保持显式。这使我们能够在序列长度上实现线性成本,同时具备窗口化2-单纯形注意力所缺乏的全局覆盖范围。我们通过自定义Triton内核实现该机制,并将其与Kimi Delta Attention结合,构建了一个完全不使用softmax注意力的模型。在计算资源匹配的情况下,该模型在对比架构中实现了最高的平均下游准确率;在16k上下文长度下,它相比KDA混合模型提高了平均下游准确率,同时将LAMBADA困惑度从715.6降至602.6。

英文摘要

We present a linearized form of 2-simplicial attention by rewriting the trilinear score as an inner product between a composite query and a key, so that the sum over one token axis takes the same form as ordinary softmax attention. We then approximate this sum with positive random features and store the entire past in a fixed-size state, while the second axis stays explicit over a short window of recent tokens. This enables us to achieve linear cost in sequence length combined with a global reach that windowed 2-simplicial attention lacks. We implement it with custom Triton kernels and combine it with Kimi Delta Attention to build a model with no softmax attention at all. Under matched compute, this model achieves the highest mean downstream accuracy among the compared architectures, and at 16k context it improves mean accuracy over a KDA hybrid while lowering LAMBADA perplexity from 715.6 to 602.6.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑