言语过程监督可培育更优的编码智能体
Verbal Process Supervision Elicits Better Coding Agents
- Mindify AI
- University of London(伦敦大学)
- National Taiwan University of Science and Technology(台湾科技大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对现有编码智能体难应对复杂软件工程任务的问题,本研究提出经言语过程监督增强的CURA系统,在BigCodeBench上较基线提升3.65%,搭配o3-mini达SOTA,推动了推理架构与LLM代码生成的融合。
AI中文摘要:
大语言模型的兴起及其作为AI智能体的应用,显著推进了代码生成基准的当前最优水平,变革了现代软件工程任务。然而,即便采用测试时计算推理模型,这类系统仍难以应对复杂软件工程挑战。本研究提出CURA——一种经言语过程监督(VPS)增强的代码理解与推理智能体系统,在BigCodeBench等具有挑战性的基准上,较基线模型实现3.65%的性能提升。此外,CURA与o3-mini模型及VPS技术结合后,达到了当前最优性能。该工作推动了推理驱动架构与基于大语言模型的代码生成的融合,使语言模型的智能体推理得以用于解决复杂软件工程任务。
英文摘要:
The emergence of large language models and their applications as AI agents have significantly advanced state-of-the-art code generation benchmarks, transforming modern software engineering tasks. However, even with test-time computed reasoning models, these systems still struggle with complex software engineering challenges. This work introduces CURA, a code understanding and reasoning agent system enhanced with verbal process supervision (VPS), achieving a 3.65\% improvement over baseline models on challenging benchmarks like BigCodeBench. Furthermore, CURA, when paired with the o3-mini model and VPS techniques, attains state-of-the-art performance. This work represents a step forward in integrating reasoning-driven architectures with LLM-based code generation, enabling agentic reasoning for language models to solve complex software engineering tasks.