AI 中文总结
本文提出GPU通信编程综合基准CommBench,评估发现最强代码生成模型GPT-5.5仅在30.7%任务中实现正确高效的GPU通信代码,凸显LLMs与专家水平的差距。
AI 中文摘要
训练和部署大语言模型(LLMs)高度依赖高性能GPU通信,但实现高效的GPU通信原语需要深入掌握GPU架构、网络硬件及分布式通信模式,这对代码生成模型而言极具挑战性。本文提出CommBench,这是一个用于GPU通信编程的综合基准,包含100余项由专家精心设计的任务,涵盖点对点通信、集合操作、专家并行通信、计算-通信融合及通信实用函数,参考实现由GPU通信专家编写或从生产代码库中提炼。我们还引入了抗作弊评估框架,可在多GPU系统上自动编译、执行并验证生成的代码,以及联合衡量功能正确性与通信性能的统一指标。在节点内NVLink和节点间RDMA平台上评估前沿及开源代码生成模型的结果显示,即便最强模型GPT-5.5也仅在30.7%的基准任务中实现正确实现并达到有竞争力的性能。我们的结果揭示了当前LLMs与专家编写的GPU通信代码间存在巨大差距,确立了CommBench作为推进AI辅助系统编程的具有挑战性的基准。
英文摘要
Training and serving large language models (LLMs) rely heavily on high-performance GPU communication, yet implementing efficient GPU communication primitives requires deep expertise in GPU architectures, networking hardware, and distributed communication patterns, making them particularly challenging for code generation models. We present CommBench, a comprehensive benchmark for GPU communication programming, consisting of over 100 expert-curated tasks spanning point-to-point communication, collective operations, expert-parallel communication, compute--communication fusion, and communication utility functions, with reference implementations either written by GPU communication experts or distilled from production codebases. We further introduce a cheat-resistant evaluation framework that automatically compiles, executes, and validates generated code on multi-GPU systems, and a unified metric that jointly measures functional correctness and communication performance. Evaluating leading frontier and open-source code generation models on both intra-node NVLink and inter-node RDMA platforms reveals that even the strongest model, GPT-5.5, correctly implements and achieves competitive performance on only 30.7\% of the benchmark tasks. Our results expose a substantial gap between current LLMs and expert-written GPU communication code, establishing CommBench as a challenging benchmark for advancing AI-assisted systems programming.