编码时自言自语:什么使注释有助于代码生成?
Talking to Itself While Coding: What Makes Comments Help Code Generation?
- Monash University(莫纳什大学)
- Singapore Management University(新加坡管理大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本研究通过观察和干预实验发现,注释对代码生成的帮助源于其传达的正确解决方案内容,而非注释本身,且提示无法可靠替代这种增益。
中文摘要 AI 辅助
大型语言模型(LLMs)在编写代码时经常生成自然语言注释,这些注释成为生成后续代码所用上下文的一部分。然而,目前尚不清楚注释的哪些属性会影响代码生成性能。我们通过观察性分析和受控干预来研究这一问题。在LiveCodeBench上,注释频率和宽泛的注释意图都不能可靠地预测pass@1。随后,我们用较强源模型编写的注释块预填充较弱的接收模型,从而将注释的表面形式与它们传达的解决方案内容分离开来。来自通过测试的源解决方案的注释平均将接收模型的pass@1提高了17.2%。相比之下,描述失败解决方案的注释没有提供可靠的增益,而为不同问题编写的注释使pass@1降低了20.8%。最后,在广泛的模型和提示变体中,大多数接收模型未能显著恢复外部注释带来的增益,最佳情况仅恢复了24%。这些结果表明,注释有助于代码生成,不仅仅因为它们是注释,而是因为它们能提供提示无法可靠引出的正确解决方案内容。
英文摘要
Large Language Models (LLMs) often generate natural-language comments while writing code, and these comments become part of the context used to generate the code that follows. However, it remains unclear which properties of comments affect code-generation performance. We study this question through observational analyses and controlled interventions. On LiveCodeBench, neither comment frequency nor broad comment intent reliably predicts pass@1. We then prefill weaker recipient models with comment blocks written by stronger source models, allowing us to separate comment surface form from the solution content they convey. Comments from source solutions that pass the tests raise recipient pass@1 by 17.2% on average. In contrast, comments describing failed solutions provide no reliable gain, while comments written for a different problem reduce pass@1 by 20.8%. Finally, across a wide range of models and prompt variants, most recipient models show no significant recovery of the external-comment gain, and the best case recovers only 24%. These results show that comments help code generation not merely because they are comments, but because they can provide correct solution content that prompting cannot reliably elicit.