arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00425cs.SEcs.AI

代码能运行,环境却不行:衡量AI生成软件的环境可复现性

Code That Works, Environments That Don't: Measuring Environment Reproducibility in AI-Generated Software

Bhanu Prakash Vangala, Tanu Malik

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出环境规范协议与三层依赖框架,评估编码智能体在环境依赖指定上的系统性错误,发现其一致性低至7%,确立环境规范为代码生成质量的独立衡量维度。

中文摘要 AI 辅助

代码生成已成为大型语言模型的核心能力,编码智能体现在能够根据自然语言提示生成功能正确的软件项目。然而,功能正确性本身并不能涵盖生成质量的一个关键维度:环境规范,即准确识别执行生成代码所需的依赖关系,同样至关重要。我们开发了一种用于环境规范的智能体协议,并引入了一个三层框架,包括声明的、运行时安装的以及必要且充分的依赖关系,以系统性地评估编码智能体的环境规范能力。使用该协议,我们评估了编码智能体在多大程度上系统性地错误指定软件环境依赖关系,以及这种错误指定在三个智能体、四种语言和五十个编程任务中的变化情况。我们的结果表明,当前的编码智能体在这一维度上表现出系统性的泛化失败,产生的依赖规范不一致、冗余或不完整,而这些问题是功能测试无法检测到的。在所有智能体中,对于相同任务,依赖集的一致性低至7%,且较新的智能体没有显示出有意义的改进,这表明该问题并未通过规模或新颖性得到解决。最大的分歧出现在声明的依赖层和运行时依赖层之间,这暗示了从模型训练分布中学习到的环境先验是主要驱动因素。我们的研究结果将环境规范确立为代码生成质量的一个独立、可测量的轴,而当前的基准并未捕捉到这一点,并激励了同时优化功能正确性和环境可移植性的训练目标和评估协议。

英文摘要

Code generation has emerged as a central capability of large language models, with coding agents now able to produce functionally correct software projects from natural language prompts. However, functional correctness alone does not capture a critical dimension of generation quality: environment specification, defined as the accurate identification of the dependencies required to execute generated code, is equally critical. We develop an agent protocol for environment specification and introduce a three-layer framework comprising declared, runtime-installed, and necessary-and-sufficient dependencies to systematically assess coding agents for environment specification. Using this protocol, we evaluate the extent to which coding agents systematically misspecify software environment dependencies and how this misspecification varies across three agents, four languages, and fifty programming tasks. Our results show that current coding agents exhibit systematic generalization failures along this dimension, producing dependency specifications that are inconsistent, redundant, or incomplete in ways that functional tests do not detect. Across agents, dependency set agreement is as low as 7% for identical tasks, and newer agents show no meaningful improvement, suggesting the failure is not resolved by scale or recency. The largest divergence occurs between the declared and runtime dependency layers, implicating environment priors learned from the models' training distributions as the primary driver. Our findings establish environment specification as a distinct, measurable axis of code generation quality that current benchmarks do not capture, and motivate training objectives and evaluation protocols that jointly optimize for functional correctness and environmental portability.

发表机构

  • University of Missouri–Columbia(密苏里大学哥伦比亚分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑