arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.01821cs.DCcs.AR

采用高基数光子互连扩展推理预填充

Scaling Inference Prefill with High-Radix Photonic Interconnects

Arulselvan Madhavan, Peter Carson, Taylor Groves, Thomas Graham

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对MoE模型的推理预填充,对比铜基GPU系统,验证3D集成光子互连可在高批处理、通信受限场景及GPU规模受限情况下,分别带来2.1--5.8倍延迟降低与2.2--4.5倍加速,满足大模型扩展需求。

中文摘要 AI 辅助

随着推理成为当今主导的AI工作负载,行业正转向高带宽光子互连,以满足日益复杂的混合专家(MoE)模型的大规模扩展需求。本文通过分析大型语言模型(LLM)聊天的高并发吞吐量与推理及智能体AI通常所需的大上下文窗口之间的权衡,量化了3D集成光子互连对推理预填充的益处。我们模拟了三种MoE模型:短上下文(1K--8K tokens)、中等上下文(128K tokens)和长上下文(1M tokens)。我们在现有基于铜的GPU系统和采用高带宽集成光子学的系统上评估该工作负载。结果显示,在高批处理压力场景下延迟提升2.1--3.2倍,在通信受限配置中较基线提升2.8--5.8倍。当电气系统达到其固有的扩展Pod限制时,3D光子学可实现1152-GPU的部署规模以降低首token生成时间,在生产级平台上实现2.2--4.5倍的加速。

英文摘要

With the rise of inference as today's dominant AI workload, the industry is transitioning to high-bandwidth photonic interconnects to meet the large scale-up requirements of increasingly complex Mixture-of-Experts (MoE) models. This paper quantifies the benefits of 3D-integrated photonic interconnects for inference prefill by analyzing tradeoffs between high-concurrency throughput for Large Language Model (LLM) chat and the large context windows typically required for reasoning and agentic AI. We simulate three MoE models: short context (1K--8K tokens), medium context (128K tokens), and long context (1M tokens). We evaluate this workload across existing copper-based GPU systems and one with high bandwidth integrated photonics. We show 2.1--3.2x latency improvements in the stressed high-batch regimes and 2.8--5.8x improvements over baselines in communication-limited configurations. 3D photonics enable the 1152-GPU footprint required to lower time-to-first-token, yielding 2.2--4.5x speedups across production-grade platforms when electrical systems cross their inherent scale-up-pod limits.

补充信息

↑