arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30288cs.CLcs.LG

流形投影与迭代自编码器精化用于掩码语言建模

Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling

Narges Mokhtari, Farzan Haddadi, Ebrahim Rezaii

首次发表
浏览论文内容

中文总结 AI 辅助

提出用低秩瓶颈自编码器堆叠替代注意力实现上下文混合,结合迭代拉取与校正精化,在掩码语言建模中以约1.9倍更少FLOPs接近注意力性能,并匹配BERT基线。

中文摘要 AI 辅助

在基于Transformer的掩码语言模型中,注意力是上下文混合的主要机制,但还有其他方式可以在token之间混合数据。最近的无注意力混合器用固定或超网络生成的MLP替换注意力,交替使用其动态的、内容相关的权重以换取计算简单性。我们构建了一种替代方案,从低秩瓶颈自编码器获得相同属性。我们用一叠基于自编码器的混合模块替换注意力,一个在局部邻域上操作,一个在整个序列上操作,一个跨注意力头操作,每个模块通过瓶颈压缩和重构其输入,其宽度是超参数而非训练效果。在掩码位置,我们引入了一个具有两个不同步骤的迭代精化过程。一个拉取步骤,将嵌入表示拉向其邻居的加权平均,以及一个校正步骤,通过自编码器将结果投影回学习到的流形。我们的架构在C4上预训练并以参数匹配的BERT基线评估时,以约$1.9 \times$更少的FLOPs实现了注意力性能的显著部分。我们的模型在稀有token频率桶上等于参数匹配的BERT和TinyBERT基线,使用频率感知的训练计划,该计划在掩码任务中对稀有token的采样比均匀采样更多。

英文摘要

In Transformer-based masked language models, attention is the primary mechanism for context mixing, but there are other ways to mix data across tokens. Recent attention-free mixers replace attention with fixed or hypernetwork-generated MLPs, alternating their dynamic, content-dependent weighting for computational simplicity. We build an alternative that gets the same property from a low-rank bottleneck autoencoder. We replace attention with a stack of autoencoder-based mixing modules, one operating over local neighborhoods, one over the full sequence, and one across attention heads, each compressing and reconstructing its input through a bottleneck, and its width is a hyperparameter rather than a training effect. In masked positions, we introduce an iterative refinement procedure that has two distinct steps. A pulling step that pulls an embedding representation toward a weighted average of its neighbors, and a correcting step that projects the result back to the learned manifold via an autoencoder. Our architecture achieves a significant portion of attention's performance at about $1.9 \times$ fewer FLOPs when pretrained on C4 and evaluated with parameter-matched BERT baselines. Our model equals parameter-matched BERT and TinyBERT baselines on the rarest-token frequency bucket using a frequency-aware training schedule that samples rare tokens more than uniformly for the masking tasks.

发表机构

  • Iran University of Science & Technology(伊朗科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑