arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2503.18494cs.AIcs.CLcs.LG

言语过程监督可培育更优的编码智能体

Verbal Process Supervision Elicits Better Coding Agents

  • Mindify AI
  • University of London(伦敦大学)
  • National Taiwan University of Science and Technology(台湾科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Hao-Yuan Chen, Cheng-Pong Huang, Jui-Ming Yao

更新

AI总结:

针对现有编码智能体难应对复杂软件工程任务的问题,本研究提出经言语过程监督增强的CURA系统,在BigCodeBench上较基线提升3.65%,搭配o3-mini达SOTA,推动了推理架构与LLM代码生成的融合。

AI中文摘要:

大语言模型的兴起及其作为AI智能体的应用,显著推进了代码生成基准的当前最优水平,变革了现代软件工程任务。然而,即便采用测试时计算推理模型,这类系统仍难以应对复杂软件工程挑战。本研究提出CURA——一种经言语过程监督(VPS)增强的代码理解与推理智能体系统,在BigCodeBench等具有挑战性的基准上,较基线模型实现3.65%的性能提升。此外,CURA与o3-mini模型及VPS技术结合后,达到了当前最优性能。该工作推动了推理驱动架构与基于大语言模型的代码生成的融合,使语言模型的智能体推理得以用于解决复杂软件工程任务。

英文摘要:

The emergence of large language models and their applications as AI agents have significantly advanced state-of-the-art code generation benchmarks, transforming modern software engineering tasks. However, even with test-time computed reasoning models, these systems still struggle with complex software engineering challenges. This work introduces CURA, a code understanding and reasoning agent system enhanced with verbal process supervision (VPS), achieving a 3.65\% improvement over baseline models on challenging benchmarks like BigCodeBench. Furthermore, CURA, when paired with the o3-mini model and VPS techniques, attains state-of-the-art performance. This work represents a step forward in integrating reasoning-driven architectures with LLM-based code generation, enabling agentic reasoning for language models to solve complex software engineering tasks.

↑