arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24637cs.AR

面向大语言模型混合专家(MoE)训练的晶圆级光互连中的热调谐开销:跨层分析与基于铁电材料的缓解方案

Cross-Layer Analysis of Thermal Tuning Stalls in Wafer-Scale Optical Interconnects for LLM MoE Training

  • Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Seongwon Yoon, Pin-Jun Chen, Shimeng Yu

AI总结:

本研究针对LLM MoE训练的晶圆级光互连,通过跨层分析发现热波动会引发调谐停滞,采用铁电电光调谐机制消除停滞,使三款MoE模型获2.7-3.8倍加速。

AI中文摘要:

大语言模型(LLM)的快速扩展,尤其是混合专家(MoE)架构,加剧了互连需求,因为专家并行执行具有高通信密集度。基于密集波分复用(DWDM)的晶圆级光互连为实现更高带宽提供了有前景的路径;然而,传统基于微环谐振器(MRR)的链路依赖热光调谐,因此易受工作负载引发的热波动影响。本研究针对MoE工作负载的晶圆级光互连开展跨层分析,结合了工作负载分析、分组级网络仿真和瞬态热分析。我们在ht-sim仿真器中实现了晶圆级拓扑,并构建了3D集成GPU/EIC/PIC堆叠的Ansys热模型。结果表明,瞬态温度变化会超出传统热光控制环路的跟踪能力,从而在通信阶段引入反复的调谐停滞。注入网络仿真的停滞时长直接源自热模型,而非假设得出。我们进一步评估了一种基于铁电材料的电光调谐机制,该机制消除了持续的热调谐需求。在针对三个MoE模型的四层代理仿真中,与热光方案相比,消除调谐停滞为Mixtral 8x7B带来2.7倍加速,为Qwen-MoE 14.3B带来3.8倍加速,为LLaMA-MoE 6.7B带来3.3倍加速。这些结果表明,最小化光子调谐延迟对发挥大规模AI系统中光互连的性能潜力至关重要。

英文摘要:

Mixture-of-experts (MoE) training depends heavily on all-to-all communication, which wafer-scale optical interconnects with dense wavelength-division multiplexing can serve. Their microring resonators rely on thermo-optic tuning, while MoE compute bursts swing the photonic-layer temperature by about 10 K within tens of milliseconds. This paper quantifies the communication stall by coupling packet-level network simulation, transient thermal simulation of a 3D-stacked GPU and photonic die, and a ring detuning criterion, and by feeding the stall back into the network timeline until both agree. Including this feedback, a tracking loop at the measured 5 nm/s lengthens the training iteration of Mixtral 8x7B and LLaMA-MoE 6.7B by factors of about 1.16 and 1.25 at full model depth. Measured H100 die temperatures match the modeled swing of 300 ms bursts within 15%. A loop slewing eight times faster, a heater driven at each kernel launch, or a dummy load near 75% of peak power removes the stall. An athermalized lithium-niobate ring with a non-volatile ferroelectric setpoint removes it with no holding power or fast loop, which makes it suitable for wafer-scale optical interconnects.

↑