arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14530cs.LGcs.CL

xHC:扩展超连接

xHC: Expanded Hyper-Connections

Xiangdong Zhang, Xiaohan Qin, Sunan Zou, Tuo Dai, Xiaoming Shi, Huaijin Wu, Yebin Yang, Zhuo Xia, Shaofeng Zhang, Lin Yao, Yuliang Liu, Yu Cheng, Junchi Yan

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对超连接(HC)扩展残差流时的性能瓶颈,提出xHC方法,结合时间特征增强与稀疏残差流架构,实现超越N = 4的有效扩展,在MoE模型上有下游改进,还介绍xHC - Flash减少内存流量,让大N残差流扩展用于语言模型预训练更有效实用。

中文摘要 AI 辅助

超连接(HC)将Transformer的残差流扩展为N个并行流,提供了一种超越模型宽度和深度的内存扩展形式。流形约束HC(mHC)在规模上稳定了这种公式。从N = 1到N = 4的巨大收益表明残差流扩展是一个有前途的扩展轴。然而,现有的HC家族方法通常在N = 4时停止。实验揭示了原因:超过此点扩展mHC会导致性能提升递减和训练成本迅速增加。将此限制归因于两个瓶颈:流数量增加时回写信息不足以及残差混合生成成本与N呈三次方缩放。为解决这两个瓶颈,提出xHC,它结合了时间特征增强以实现更丰富的回写,并采用稀疏残差流架构,仅更新N = 16个流中的k = 4个流,同时保留对完整残差状态的密集访问。在跨18B和28B的混合专家(MoE)模型上,xHC实现了强大且一致的下游改进。还介绍了xHC - Flash,它减少了子层内存流量,使大N残差流扩展对语言模型预训练有效且实用。

英文摘要

Hyper-Connections (HC) expand the residual stream of Transformers into $N$ parallel streams, providing a form of memory scaling beyond model width and depth. Manifold-Constrained HC (mHC) stabilizes this formulation at scale. The large gains from $N{=}1$ to $N{=}4$ suggest residual-stream expansion as a promising scaling axis. However, existing HC-family methods typically stop at $N{=}4$. Our experiments reveal why: scaling mHC beyond this point yields diminishing performance gains and rapidly increasing training cost. We attribute this limitation to two bottlenecks: insufficient write-back information for an expanding number of streams and residual-mixing generation whose cost scales cubically with $N$. To address both bottlenecks, we propose xHC (Expanded Hyper-Connections), the first HC-family method to achieve meaningful expansion beyond $N{=}4$. xHC combines temporal feature augmentation for richer write-back with a sparse residual-stream architecture that updates only $k=4$ of the $N=16$ streams while retaining dense access to the full residual state. Across 18B and 28B MoE models, xHC delivers strong and consistent downstream improvements. On an 18B MoE model, xHC improves the average downstream score by 4.0 points over mHC, while adding only modest training FLOPs over the vanilla baseline. Scaling-law experiments show that the vanilla and mHC require $1.50\times$ and $1.19\times$ the compute of xHC, respectively, to reach the same loss. Practical large-$N$ training also requires controlling memory traffic from the expanded residual state. We therefore introduce xHC-Flash, which reduces the per-sublayer memory traffic from $73.5C$ to $40C$, comparable to the $34C$ required by mHC at $N{=}4$, while retaining the gains of full xHC. Together, xHC and xHC-Flash make large-$N$ residual-stream expansion effective and practical for LLM pre-training.

发表机构

  • School of AI, Shanghai Jiao Tong University(上海交通大学人工智能学院)
  • Dots Studio, Xiaohongshu Inc.(小红书公司点点工作室)
  • University of Science and Technology of China(中国科学技术大学)
  • School of CS, Peking University(北京大学计算机科学学院)
  • The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑