arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高效Transformer中的稀疏令牌路由

Sparse Token Routing in Efficient Transformers

Sai Krishna Arthanari, JaeHyeong Chang, Chengzhe Sun, Siwei Lyu

arXiv 2608.20632首次发表:更新:

发表机构

Institute for Artificial Intelligence and Data Science (IAD); University at Buffalo(人工智能与数据科学研究所; 布法罗大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对高效Transformer的令牌计算分配问题,采用SEWN模型的双流路由机制,验证了令牌路由对精度影响极小,且完全上下文门控在两项任务上实现了高度显著分离,为高效Transformer的稀疏令牌路由提供了有效方案。

AI 中文摘要

高效Transformer研究常以并非所有令牌需要同等计算量为依据,推动令牌剪枝与自适应计算。我们采用SEWN(一种双流Transformer,通过学习门控将令牌路由至轻量或全容量处理)端到端验证该主张。实验显示,与参数匹配的基线相比,路由引入的精度变化可忽略不计,而门控的令牌重要性信号关键取决于其学习方式。基于静态词典的先验在BoolQ上的反事实忠实性测试中失败,而完全上下文门控在两项评估任务上实现了高度显著的分离(p<10^-10),且未改变任务精度。

英文摘要

Efficient-transformer research often motivates token pruning and adaptive computation with the claim that not all tokens require equal computational effort. We test this claim end to end using SEWN, a two-stream Transformer that routes tokens through either lightweight or full-capacity processing using a learned gate. Across our experiments, routing introduces negligible accuracy change relative to parameter-matched baselines, while the gate's token-importance signal depends critically on how it is learned. A static lexicon-seeded prior fails a counterfactual faithfulness test on BoolQ, whereas a fully contextual gate achieves highly significant separation ($p<10^{-10}$) on both evaluated tasks without changing task accuracy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑