arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ParBench:用于可靠评估大语言模型并行代码翻译的基准测试

ParBench: A Benchmark for Reliable Evaluation of LLM Parallel Code Translation

Samyak Jhaveri, Erel Kaplan, Tom Yotam, Le Chen, Tomer Bitan, Niranjan Hasabnis, Gal Oren

arXiv 2607.22588首次发表:更新:

AI 中文总结

研究针对大语言模型并行代码翻译缺乏可靠评估方法的问题,提出ParBench基准框架,通过声明性规范评估,涵盖多开源套件与跨API翻译方向,经源增强测试,揭示了当前并行代码翻译存在的如方向不对称等障碍。

AI 中文摘要

现代计算密集型软件需在不断变化的加速器、编程API、编译器堆栈和可移植层生态系统中迁移,大语言模型和自主编码代理被用于此迁移,但缺乏可靠方法衡量其是否保留使翻译行为有效的低级并行语义。本文提出ParBench,这是一个以内核为中心的基准框架,用于在可执行、可重现条件下评估基于大语言模型的并行API翻译。它通过声明性基准规范修复周边构建、运行和验证基础设施,仅要求模型翻译计算内核,涵盖多个开源HPC套件及代表性跨API翻译方向。还包括AST驱动、保留预期行为、经基线验证的源增强以测试成功是否反映稳健翻译而非表面形式记忆。对现有大语言模型的评估显示了可靠并行代码翻译存在持续障碍。

英文摘要

Modern compute-intensive software must migrate across a changing ecosystem of accelerators, programming APIs, compiler stacks, and portability layers, including CUDA, OpenMP, OpenCL, and OpenMP target offload. Large language models and autonomous coding agents are increasingly proposed for such migration, but the field lacks reliable ways to measure whether they preserve the low-level parallel semantics that make translations behaviorally valid, including thread indexing, synchronization, memory management, host-device coordination, and API-specific execution structure. We present ParBench, a kernel-centric benchmark framework for evaluating LLM-based parallel API translation under executable, reproducible conditions. ParBench fixes the surrounding build, run, and verification infrastructure through declarative benchmark specifications and asks models to translate only the computational kernels. It draws on multiple open-source HPC suites and covers representative cross-API translation directions among CUDA, OpenMP, OpenCL, and OpenMP target offload. To test whether success reflects robust translation rather than surface-form memorization, ParBench includes AST-driven, intended behavior-preserving, baseline-validated source augmentation. Evaluations on state-of-the-art open and proprietary LLMs show persistent barriers to reliable parallel code translation, including direction asymmetry, multi-file coordination, incomplete API adaptation, and uneven robustness to source-level perturbations. Code is available at https://github.com/Scientific-Computing-Lab/ParBench.

Comments57 pages, 22 figures, 11 tables; includes appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑