arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18311cs.SE

提示词中自然语言差异对基于LLM的自动代码生成影响的研究

A Study on the Impact of Natural Language Differences in Prompts on Automatic Code Generation Using LLMs

Haruka Tokumasu, Masanari Kondo, Alexander Serebrenik, Dong Wang, Kei Koyanagi, Kotaro Noguchi, Naoyasu Ubayashi, Yasutaka Kamei

首次发表
浏览论文内容

中文总结 AI 辅助

本研究量化了提示词自然语言对LLM代码生成准确性的影响,发现语言偏差显著,翻译可部分缓解但效果因数据集和模型而异。

中文摘要 AI 辅助

大型语言模型(LLMs)在自动代码生成任务中展现了卓越的性能,从而促进了该领域的新研究。尽管已有大量研究探索了基于LLM的代码生成,但输入提示词中自然语言的影响(语言偏差)仍未得到充分探索。本研究旨在(1)量化输入提示词的自然语言如何影响基于LLM的代码生成性能,以及(2)评估一种减少代码生成中语言偏差的缓解策略。我们在AtCoder、LeetCode和BigCodeBench上评估代码生成的准确性。为了量化代码生成中的语言偏差,每个问题分别以英语、日语和中文呈现。我们使用了七个LLM(GPT-4o、o3-mini、DeepSeek-V3.2、Llama-3、Qwen2.5-Coder-14B、Qwen2.5-Coder-0.5B和GitHub Copilot),并以其准确性(生成的代码通过所有测试用例的问题数量)来评估其性能。我们比较了翻译前后的准确性,以评估翻译作为缓解策略的有效性。我们观察到,问题陈述的自然语言影响了基于LLM的代码生成性能。具体来说,每个数据集官方支持的语言取得了最高的中位准确性。此外,翻译提高了准确性,但其有效性在不同数据集和模型类型之间并不一致。我们发现AtCoder包含特别高比例的叙事风格问题陈述和较长的问题陈述。自然语言显著影响LLM代码生成的准确性。翻译可以在某些设置下缓解语言偏差,但其有效性取决于数据集和模型类型。此外,输入提示词的叙事方面和上下文长度是与语言偏差及翻译作为缓解策略有效性相关的重要因素。

英文摘要

Large Language Models (LLMs) have demonstrated remarkable performance in automatic code generation tasks, thereby encouraging new research in this area. Although numerous studies have explored LLM-based code generation, the impact of the natural language in input prompts remains unexplored (language bias). This study aims to (1) quantify how the natural language of input prompts influences LLM-based code generation performance and (2) evaluate a mitigation strategy to reduce language bias in code generation. We assess code generation Accuracy on AtCoder, LeetCode, and BigCodeBench. To quantify the language bias on code generation, each problem is presented in English, Japanese, and Chinese. We use seven LLMs (GPT-4o, o3-mini, DeepSeek-V3.2, Llama-3, Qwen2.5-Coder-14B, Qwen2.5-Coder-0.5B, and GitHub Copilot) and assess their performance in terms of Accuracy (the number of problems for which generated code passes all test cases). We compare Accuracy before and after translation to evaluate the effectiveness of translation as a mitigation strategy. We observed that the natural language of problem statements affects LLM-based code generation performance. Specifically, the languages officially supported by each dataset achieved the highest median Accuracy. Also, translation improved Accuracy, but its effectiveness was not consistent across datasets and model types. We found that AtCoder contained a particularly high proportion of narrative-style problem statements and longer problem statements. Natural language significantly affects LLM code generation accuracy. Translation can mitigate language bias in some settings, but its effectiveness depends on the dataset and model type. Furthermore, the narrative aspects and context length of input prompts are important factors related to language bias and the effectiveness of translation as a mitigation strategy.

发表机构

  • Kyushu University(九州大学)
  • Eindhoven University of Technology(埃因霍温理工大学)
  • Tianjin University(天津大学)
  • Waseda University(早稻田大学)
  • Inamori Research Institute for Science(稻盛科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑