发表机构
City University of Hong Kong; Jilin University; National University of Singapore(香港城市大学; 吉林大学; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出Relation这一token混合原语,衍生出多种关系变体,在不同规模模型上验证其性能,FlashRelation速度优势明显,Hybrid Relation兼具效率与质量,为token混合提供关系优先思路。
AI 中文摘要
注意力机制直接从成对分数中推导归一化信息流。我们引入了Relation(关系)这一替代的token混合原语,它先将成对证据组织为显式的Self(自身)与Exchange(交互)关系,之后再推导信息流。这种关系组织方式衍生出了Full Relation(全关系)、FlashRelation(快速关系)、Linear Relation(线性关系)、Hybrid Relation(混合关系)以及一种KV风格的Relation Cache(关系缓存)。在参数规模约为10M、30M和100M的匹配解码器仅模型中,Full Relation在所有三个规模下都取得了比MHA(多头注意力)更低的最终验证NLL(负对数似然)。在固定上下文参考基准中,FlashRelation的运行速度是具体化的Full Relation实现的3.60-4.41倍。在规模匹配的生产工作负载中,它在执行Full Relation算子时达到了PyTorch FlashAttention吞吐量的76.4-84.9%。Hybrid Relation使用75%的Linear Relation层,且具备出色的语言建模质量。这些结果支持了token混合的关系优先观点:询问自身,询问他者,随后让信息流跟随关系。
英文摘要
Attention dominates token mixing, but it collapses relation formation and flow allocation into a single score-to-flow step. We introduce Relation, which separates them by first organizing pairwise evidence into explicit Self and Exchange relations and deriving information flow afterward. Relation first decides whether a token should rely on itself or draw from its history, and if it draws from history, where to look. This relational organization gives rise to Full Relation, FlashRelation, Linear Relation, and Hybrid Relation. Across matched decoder-only models, Full Relation achieves lower mean final-validation NLL than MHA and reaches the paired MHA final training loss with 4.5-7.3% fewer tokens. Structural diagnostics further show that Relation learns a distinct depth organization: the first layer acts as a current-token anchor and a high-rank router, while later layers shift strongly toward history. In a fixed-context reference benchmark, FlashRelation is 4.17-5.28x faster than the materialized Full Relation implementation. Across scale-matched production workloads, it reaches 89.7-92.9% of PyTorch FlashAttention throughput while executing the exact Full Relation operator. Hybrid Relation demonstrates that Full and Linear Relation layers can be composed within a single decoder. These results support a relation-first view of token mixing: ask Self, ask Others, then let Flow follow Relation.