arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于从多语言提示生成代码的大语言模型:一个精心策划的基准测试和对代码质量的研究

Large Language Models for Code Generation from Multilingual Prompts: A Curated Benchmark and a Study on Code Quality

Saima Afrin, Alessandro Midolo, Camilo Escobar-Velásquez, Mario Linares-Vásquez, Weiyuan Ding, Bowen Xu, Massimiliano Di Penta, Antonio Mastropaolo

arXiv 2607.14816首次发表:更新:

发表机构

University of Catania; North Carolina State University; University of Sannio(卡塔尼亚大学; 北卡罗来纳州立大学; 萨尼奥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究多语言提示对代码生成的影响,通过460个Python和Java编码任务,将英文提示翻译成多种语言并评估生成代码,发现英文提示非最佳,提示语言影响因编程语言和大语言模型而异,提供多语言基准测试及见解。

AI 中文摘要

大语言模型(LLMs)在相同编程任务中使用不同自然语言提示时表现不同,即语言偏差现象。虽该行为在一般文本生成中被广泛研究,但对代码生成质量和编程惯例的影响仍未充分探索。我们研究描述编程任务的语言如何影响GPT-4o mini、DeepSeek和Claude生成的源代码。研究包含460个Python和Java编码任务。我们将原始英文提示翻译成中文、印地语、西班牙语和意大利语并手动策划,同时保留技术含义。我们从多个维度评估生成的代码。结果表明:英文提示不一定产生最佳功能正确性或代码质量;提示语言的影响取决于编程语言和大语言模型;生成的代码在注释和字符串字面量中经常将英文与提示语言混合。这些发现为研究代码生成中的语言偏差提供了首个精心策划的多语言基准测试,并为开发更强大的多语言代码生成系统提供了见解。

英文摘要

Large Language Models (LLMs) perform differently on identical programming tasks when prompted in different natural languages, a phenomenon known as language bias. While this behavior has been widely studied for general text generation, its impact on code generation quality and programming conventions remains largely unexplored. We investigate how the language used to describe programming tasks affects the source code generated by GPT-4o mini, DeepSeek, and Claude. Our study comprises 460 coding tasks spanning Python (230) and Java (230). We translate and manually curate the original English prompts into Chinese, Hindi, Spanish, and Italian while preserving their technical meaning. We evaluate the generated code using multiple dimensions, including functional correctness through test pass rates, structural quality using established code metrics, issues detected by static analysis tools, and lexical characteristics such as the language used in identifiers and comments. Our results show that (i) English prompts do not consistently produce the best functional correctness or code quality, (ii) the impact of prompt language depends on both the programming language and the LLM, and (iii) generated code frequently mixes English with the prompt language in comments and string literals. These findings provide the first curated multilingual benchmark for studying language bias in code generation and offer insights for developing more robust multilingual code generation systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑