arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DSEffi-Bench:揭示大语言模型在高效数据科学代码生成中的能力

DSEffi-Bench: Demystifying Large Language Models' Capability in Efficient Data Science Code Generation

Zhihao Gong, Junzhe Yu, Dong Huang, Zeyu Sun, Jie M. Zhang, Dan Hao

arXiv 2608.30248首次发表:更新:

发表机构

Key Lab of HCST (PKU), MOE; SCS, Peking University, China; School of Electronic and Computer Engineering, PKU Shenzhen Graduate School, China; National University of Singapore, Singapore; Institute of Software, Chinese Academy of Sciences, Beijing, China; King’s College London, London, United Kingdom(北京大学计算机辅助设计与图形学教育部重点实验室;北京大学信息科学技术学院; 北京大学深圳研究生院电子与计算机工程学院; 新加坡国立大学; 中国科学院软件研究所; 伦敦国王学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推出首个针对大语言模型生成数据科学代码执行效率的基准DSEffi-Bench,评估发现正确性与效率不相关,分类法可指导优化以提升效率并降低成本。

AI 中文摘要

当前的数据科学(DS)代码生成基准将正确性等同于质量,却忽略了正确解决方案之间存在数量级差异的执行时间。我们推出DSEffi-Bench,这是首个专门针对大语言模型(LLM)生成的DS代码执行效率的基准,包含10多个DS库中的1000个实例,配备压力测试工具和人工验证的参考方案。我们评估了3个层级的16个模型,发现仅正确性无法表征效率:GPT-5.4在正确性(Pass,66.9%)上领先,但其效率得分(B|P,71.7%)与解决任务数少47个的GPT-5.4-mini(71.6%)几乎持平;Kimi-K2.5在前沿模型中正确性最低(40.2%),但在所有16个模型中效率得分最高(73.6%)。人工标注的五类分类法显示,79.1%的效率缺陷超出算法复杂度范畴,源于特定领域的根本原因,且不同模型层级和库存在不同的失败特征。两项探索性实验提供初步证据,这些诊断可指导改进,通过分类法引导优化可实现最高14.7%的效率提升,通过库条件路由可在成本低13.0倍的情况下达到Claude-Opus-4.6 Best@3的效率。

英文摘要

Current data science (DS) code generation benchmarks equate correctness with quality, overlooking execution time differences that span orders of magnitude between correct solutions. We introduce DSEffi-Bench, the first benchmark specifically targeting execution efficiency in LLM-generated DS code, comprising 1,000 instances across 10+ DS libraries with stress-testing harnesses and human-validated references. Evaluating 16 models across 3 tiers, we find that correctness alone fails to characterize efficiency: GPT-5.4 leads in correctness (Pass, 66.9\%) but its efficiency score (B$|$P, 71.7\%) nearly matches GPT-5.4-mini (71.6\%), which solves 47 fewer tasks; Kimi-K2.5 ranks lowest in correctness among frontier models (40.2\%) yet achieves the highest efficiency score (73.6\%) across all 16 models. A human-annotated five-category taxonomy reveals that 79.1\% of efficiency deficits extend beyond algorithmic complexity to domain-specific root causes, with distinct failure profiles across model tiers and libraries. Two exploratory experiments provide initial evidence that these diagnostics can guide improvement, yielding up to +14.7\% efficiency gains via taxonomy-guided optimization and approaching Claude-Opus-4.6 Best@3 in efficiency at 13.0$\times$ lower cost via library-conditioned routing.

CommentsAccepted to EMNLP 2026 Main Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑