发表机构
The University of Hong Kong; National University of Singapore; The Chinese University of Hong Kong, Shenzhen(香港大学; 新加坡国立大学; 香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出ORCA基准评估大语言模型在数据科学代码翻译中的表现,发现现有模型性能有限,并提出意图增强方法,在ORCA-MAIN和ORCA-PROJECT上分别提升成功率4.80%和5.33%。
AI 中文摘要
数据科学代码翻译(DSCT)是将代码在数据科学库之间进行转换,同时保持功能等价性并实现数据科学生态系统互操作性的过程。尽管大语言模型(LLMs)在数据科学代码生成(DSCG)方面已展现出显著进展,但其在DSCT中的表现仍研究不足。为填补这一空白,我们提出了ORCA,一个包含两个互补设置的综合性基准:ORCA-MAIN,包含3个代表性领域(数据查询、数据操作和深度学习)中的1,600个精心策划的基础级任务;以及ORCA-PROJECT,包含覆盖7种数据科学任务类型的完整数据科学项目上的200个翻译任务。每个任务都附有带注释的参考翻译和用于验证功能等价性的测试用例。我们进一步整合了一个多阶段质量验证流程,以彻底验证任务正确性和测试用例的鲁棒性。实验结果表明DSCT面临挑战,即使是前沿LLMs也表现出有限的性能。具体而言,Claude-Opus-4.6在ORCA-MAIN上的成功率为56.92%,在ORCA-PROJECT上为33.67%,表明DSCT仍有相当大的改进空间。我们还观察到DSCT中存在明显的方向偏好,当源代码通过更明确、细粒度的操作来表达任务时,翻译始终更容易。受此启发,我们提出了一种意图增强方法,其中模型首先推断源代码意图,然后将其用作翻译的额外上下文,在ORCA-MAIN和ORCA-PROJECT上分别实现了平均绝对成功率提升4.80%和5.33%。
英文摘要
Data Science Code Translation (DSCT) is the process of converting code between data science libraries while preserving functional equivalence and enabling interoperability across data science ecosystems. While Large Language Models (LLMs) have demonstrated considerable progress in Data Science Code Generation (DSCG), their performance in DSCT remains insufficiently studied. To address this gap, we introduce ORCA, a comprehensive benchmark with two complementary settings: ORCA-MAIN, which comprises 1,600 carefully curated grounding-level tasks across 3 representative domains: Data Querying, Data Manipulation, and Deep Learning; and ORCA-PROJECT, which contains 200 translation tasks over complete data science projects across 7 data science task types. Each task is accompanied by annotated reference translations and test cases for validating functional equivalence. We further incorporate a multi-stage quality verification process that thoroughly verifies task correctness and test case robustness. Experimental results demonstrate challenges in DSCT, with even frontier LLMs showing limited performance. Specifically, Claude-Opus-4.6 achieves a success rate of 56.92% on ORCA-MAIN and 33.67% on ORCA-PROJECT, indicating considerable room for improvement in DSCT. We also observe a clear directional preference in DSCT, where translation is consistently easier when the source code expresses the task through more explicit, fine-grained operations. Motivated by this, we propose an intent-augmented method, in which the model first infers source-code intent and then uses it as additional context for translation, achieving average absolute success-rate gains of 4.80% and 5.33% on ORCA-MAIN and ORCA-PROJECT, respectively.
Comments36 pages, 15 figures, 24 tables