采用高基数光子互连扩展推理预填充
Scaling Inference Prefill with High-Radix Photonic Interconnects
浏览论文内容
中文总结 AI 辅助
该研究针对MoE模型的推理预填充,对比铜基GPU系统,验证3D集成光子互连可在高批处理、通信受限场景及GPU规模受限情况下,分别带来2.1--5.8倍延迟降低与2.2--4.5倍加速,满足大模型扩展需求。
中文摘要 AI 辅助
随着推理成为当今主导的AI工作负载,行业正转向高带宽光子互连,以满足日益复杂的混合专家(MoE)模型的大规模扩展需求。本文通过分析大型语言模型(LLM)聊天的高并发吞吐量与推理及智能体AI通常所需的大上下文窗口之间的权衡,量化了3D集成光子互连对推理预填充的益处。我们模拟了三种MoE模型:短上下文(1K--8K tokens)、中等上下文(128K tokens)和长上下文(1M tokens)。我们在现有基于铜的GPU系统和采用高带宽集成光子学的系统上评估该工作负载。结果显示,在高批处理压力场景下延迟提升2.1--3.2倍,在通信受限配置中较基线提升2.8--5.8倍。当电气系统达到其固有的扩展Pod限制时,3D光子学可实现1152-GPU的部署规模以降低首token生成时间,在生产级平台上实现2.2--4.5倍的加速。
英文摘要
With the rise of inference as today's dominant AI workload, the industry is transitioning to high-bandwidth photonic interconnects to meet the large scale-up requirements of increasingly complex Mixture-of-Experts (MoE) models. This paper quantifies the benefits of 3D-integrated photonic interconnects for inference prefill by analyzing tradeoffs between high-concurrency throughput for Large Language Model (LLM) chat and the large context windows typically required for reasoning and agentic AI. We simulate three MoE models: short context (1K--8K tokens), medium context (128K tokens), and long context (1M tokens). We evaluate this workload across existing copper-based GPU systems and one with high bandwidth integrated photonics. We show 2.1--3.2x latency improvements in the stressed high-batch regimes and 2.8--5.8x improvements over baselines in communication-limited configurations. 3D photonics enable the 1152-GPU footprint required to lower time-to-first-token, yielding 2.2--4.5x speedups across production-grade platforms when electrical systems cross their inherent scale-up-pod limits.