代码推理用于软件工程任务:调查与呼吁行动
Code Reasoning for Software Engineering Tasks: A Survey and A Call to Action
- IBM
- Columbia University(哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文调查代码推理技术,探讨其在软件工程任务中的影响,提出未来研究方向。
AI中文摘要:
大语言模型的兴起使自然语言任务取得显著进步,其在某些任务上的表现可通过结合测试时推理技术进一步提升。这些推理技术已应用于代码领域,使代码生成、测试生成和问题解决等复杂软件工程任务成为可能。然而,不同推理技术对代码导向的软件工程任务的影响尚未系统研究。本文调查了支撑这些能力的代码推理技术,重点探讨测试时计算和推理时推理范式。我们审视了多种代码特定的推理方法,并逐步构建出结合规划、工具使用和多步骤交互的软件工程代理。我们还比较了不同技术对编码任务的影响,突显其相对重要性,并概述了开放挑战和未来研究方向。在常用模型和基准测试中,利用代码特定信号(如结构和执行反馈)的方法经常与提升性能相关,这促使我们进一步研究代码推理,而不仅仅是自然语言推理。
英文摘要:
The rise of large language models (LLMs) has led to dramatic improvements across a wide range of natural language tasks. Their performance on certain tasks can be further enhanced by incorporating test-time reasoning techniques. These inference-time advances have been adopted into the code domain, enabling complex software engineering (SWE) tasks such as code generation, test generation and issue resolution. However, the impact of different reasoning techniques on code-centric SWE tasks has not been systematically explored. In this work, we survey code reasoning techniques that underpin these capabilities, with a focus on test-time compute and inference-time reasoning paradigms. We examine a variety of code-specific reasoning methods and progressively build up to SWE agents, which combine planning, tool use, and multi-step interaction. We also compare the impact of different techniques on coding tasks, highlighting their relative importance and outlining open challenges and future research directions. Across commonly used models and benchmarks, we find that approaches exploiting code-specific signals (e.g., structure and execution feedback) are frequently associated with improved performance, motivating a dedicated study of code reasoning beyond natural-language reasoning.