AI 中文总结
本文评估Cursor's Composer 2.0等三款AI编程智能体在三种语言三类算法上的并行代码生成性能,发现其并行能力弱于串行,性能提升高度依赖算法与语言,需将运行时效率纳入大语言模型核心指标。
AI 中文摘要
AI编程智能体已在软件工程领域普及,其串行性能(准确率与速度)已被广泛研究,但近期初步结果显示,它们的并行编程能力落后于串行编程能力。本文对Cursor's Composer 2.0、GPT 5.4、Claude Sonnet 4.6三款编程智能体开展跨语言评估,覆盖C++、Python、Julia三种语言,以及排序、图遍历、搜索三类算法的并行代码生成任务。针对每一组算法与语言对,研究人员提示智能体从串行基线生成并行实现,记录达到功能正确性与性能提升所需的提示工作量,并测量其相对于自定义串行基线和第三方库实现的加速比。研究发现,编程智能体可通过适度提示生成正确的并行实现,但实现显著加速比高度依赖算法与语言:Sonnet 4.6的整体性能提升最强,GPT 5.4虽能保证代码正确性但无可测量加速比;C++的图算法并行化表现最稳定,Python和Julia在搜索算法上的加速比最大,无单一语言在所有类别中占优,两者在部分图算法上能实现加速但在其他图算法上出现性能倒退。这些结果强调,除准确率外,应将运行时性能效率作为大语言模型的核心性能指标,尤其针对并行实现场景。
英文摘要
AI coding agents have quickly become omnipresent in software engineering. Their serial performance, both in terms of accuracy and speed, has been extensively covered. However, recent initial results suggest their parallel programming capabilities lag behind serial programming capabilities. This paper presents a cross-language evaluation of three coding agents -- Cursor's Composer 2.0, GPT 5.4, and Claude Sonnet 4.6 -- on parallel code generation across three algorithm categories -- sorting, graph traversal, and search -- in C++, Python, and Julia. For each algorithm and language pair, we prompt a coding agent to produce a parallel implementation from a serial baseline, track the prompting effort required to achieve both functional correctness and performance improvements, and measure speedup against both custom serial baselines and third-party library implementations. We find that coding agents can produce correct parallel implementations with modest prompting effort, but that achieving meaningful speedup is heavily algorithm- and language-dependent. Sonnet 4.6 delivers the strongest overall performance gains, whereas GPT 5.4 produces no measurable speedups despite consistent correctness. C++ is most consistently parallelizable for graph algorithms, while Python and Julia achieve the largest speedups on search algorithms: no single language dominates across all categories. Python and Julia each achieve speedup on some graph algorithms but regress on others. These findings underscore the impact of including runtime performance efficiency as a main LLM performance metric, in addition to accuracy, particularly for parallel implementations.