放置是自由的,组合则不然:拉丁方作为异构序列混合器堆栈的可证明平衡构造
Placement Is Free, Composition Is Not: The Latin Square as a Provably-Balanced Construction for Heterogeneous Sequence-Mixer Stacks
浏览论文内容
中文总结 AI 辅助
本研究通过拉丁方构造证明异构序列混合器堆栈的性能主要取决于跨深度的组合而非放置位置,并发布Aether-7B-5Attn模型及完整资源。
中文摘要 AI 辅助
自GPT以来,大多数Transformer在每一层都重复相同的注意力机制。然而,这种设计在很大程度上是一种惯例,而非经过验证的结论。当多个序列混合器组合在一个堆栈中时,性能提升可能源于机制选择、放置位置或两者兼有,这使得因果归因变得困难。我们引入了Aether-7B-5Attn,一个具有65.9亿参数的混合专家模型(约29.8亿激活参数),其49层包含七种序列混合机制,排列成一个7×7的拉丁方。由于每种机制在每一行和每一列中恰好出现一次,该设计保证了跨深度的平衡暴露,同时消除了放置混淆。为了评估这一原则,我们构建了一个参数匹配的代理模型,将四种机制排列成16层上的4×4拉丁方,匹配7.009亿参数,并为每个臂使用八个种子进行训练。结果揭示了明确的分离。将分布式的异构堆栈重新排列为平衡的周期性循环,验证损失仅变化0.16%,表明精确放置影响甚微。相比之下,将相同机制聚类到连续的深度带中会产生0.59%的惩罚,而用同质堆栈替换异构堆栈则会产生1.68%的惩罚。这些结果表明,性能主要取决于分布在深度上的异构组合,而非任何特定的排列。我们在2.16倍更大的规模(15.14亿参数)上证实了这一发现,其中同质堆栈的惩罚增加到2.63%,而移除SSM家族机制则产生3.20%的退化。我们进一步报告了每种机制的成本概况、英语和韩语评估,以及对所有49层的因果安全审计。我们发布了模型权重、训练配方、训练代码、日志和架构源代码。
英文摘要
Since GPT, most Transformers have repeated the same attention mechanism at every layer. Yet this design is largely a convention rather than a tested conclusion. When multiple sequence mixers are combined in one stack, improvements may arise from mechanism choice, placement, or both, making causal attribution difficult. We introduce Aether-7B-5Attn, a 6.59B-parameter mixture-of-experts model ($\approx$2.98B active) whose 49 layers contain seven sequence-mixing mechanisms arranged as a $7\times7$ Latin square. Because each mechanism appears exactly once in every row and column, the design guarantees balanced exposure across depth while eliminating placement confounds. To evaluate this principle, we build a parameter-matched proxy with four mechanisms arranged as a $4\times4$ Latin square over sixteen layers, matched to 700.9M parameters and trained with eight seeds per arm. The results reveal a clear dissociation. Rearranging a distributed heterogeneous stack into a balanced periodic cycle changes validation loss by only 0.16\%, indicating that exact placement has little effect. In contrast, clustering the same mechanisms into contiguous depth bands incurs a 0.59\% penalty, while replacing the heterogeneous stack with a homogeneous one incurs a 1.68\% penalty. These results indicate that performance depends primarily on heterogeneous composition distributed across depth rather than on any particular permutation. We confirm this finding at 2.16$\times$ larger scale (1.514B parameters), where the homogeneous-stack penalty increases to 2.63\% and removing the SSM-family mechanism produces a 3.20\% degradation. We further report per-mechanism cost profiles, English and Korean evaluations, and a causal-safety audit of all 49 layers. We release model weights, training recipes, training code, logs, and architecture source code.
发表机构
- VIDRAFT AI Research(VIDRAFT AI 研究院)
- QuantumOS
机构由 AI 辅助整理,请以论文原文为准。