GPUPhysBench:用于正确且高效的GPU物理模拟的编码智能体基准测试
GPUPhysBench: Benchmarking Coding Agents for Correct and Efficient GPU Physics Simulation
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
GPUPhysBench是一个包含50个任务的基准测试,用于评估编码智能体在GPU物理模拟中正确且高效地实现数值方法的能力,发现即使最强模型在效率上仍有显著差距。
AI中文摘要:
为物理模拟编写快速的GPU代码是困难的:实现必须在处理不规则数据访问、同步和迭代求解器的同时保持数值精度。我们引入了GPUPhysBench,一个包含50个任务的基准测试,用于测试编码智能体是否能够满足这些需求。任务涵盖流体、可变形固体和颗粒材料,从单个模拟算子到完整的模拟器。智能体在固定的时间预算内,在可访问NVIDIA GPU的情况下编写、编译、测试和优化GPU代码。我们报告了相对于专家优化参考实现的通过率和运行时性能。在对六个前沿模型-工具组合的单次尝试评估中,最强的两个通过了所有50个任务,但即使是最快的也仅在22%的任务上达到了至少0.9倍的参考速度,并且没有提交物比参考快超过5%。最大的差距出现在碰撞检测、约束求解和迭代求解器中。GPUPhysBench将物理模拟工作负载引入编码智能体评估,既测试了正确实现数值方法的能力,也测试了使其高效运行的能力。
英文摘要:
Writing fast GPU code for physical simulation is difficult: implementations must preserve numerical accuracy while handling irregular data access, synchronization, and iterative solvers. We introduce GPUPhysBench, a benchmark of 50 tasks testing whether coding agents can meet these demands. Tasks cover fluids, deformable solids, and granular materials, from individual simulation operators to complete simulators. Agents write, compile, test, and optimize GPU code with access to a NVIDIA GPU under fixed time budgets. We report pass rates and runtime performance relative to expert-optimized reference implementations. In a single-attempt evaluation of six frontier model-harness pairs, the two strongest pass all 50 tasks, but even the fastest reaches at least 0.9 the reference speed on only 22% of them, and no submission is more than 5% faster than the reference. The largest gaps arise in collision detection, constraint solving, and iterative solvers. GPUPhysBench brings physical simulation workloads to coding-agent evaluation, testing both the ability to implement numerical methods correctly and the ability to make them run efficiently.