arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02329cs.SE

人工智能生成代码的自动化且可解释的来源研究

On Automated and Explainable Provenance of AI-Generated Code

Alejandro Velasco, Nathan Wintersgill, Trevor Stalnaker, Oscar Chaparro, Denys Poshyvanyk

首次发表
浏览论文内容

中文总结 AI 辅助

针对AI生成代码来源不透明的问题,本文提出需构建可解释来源的下一代CodeGenAI工具,结合实证研究与因果可解释技术,明确了相关研究挑战。

中文摘要 AI 辅助

用于代码生成的生成式人工智能(CodeGenAI)已改变了软件开发,但也带来了关键的透明度问题:使用AI生成代码的开发者、部署它的组织以及负责确保其法律和质量标准的合规专业人员,都无法知晓AI生成代码的来源。现有的缓解措施仅在事后标记有问题的输出,却未解释模型为何生成这些输出,以及如何改进未来的生成。本文提出了一项基于美国国家科学基金会(NSF)资助研究项目的研究愿景,主张下一代CodeGenAI工具必须建立在可解释来源的基础上:即自动化的事后可追溯性,将生成的代码与导致其生成的提示组件、训练数据实例、全局数据特征以及模型内部组件关联起来。我们通过对软件开发者、模型用户以及合规/法律专业人员的研究获得的实证证据支撑了这一愿景,这些证据表明,来源信息是当前工具未提供的实际必要需求。我们从四个可追溯性维度对该问题进行了刻画,概述了一项结合大规模实证研究与事后因果及可解释性技术的研究计划,并确定了为实现这一愿景,研究界必须解决的关键开放挑战。

英文摘要

Generative AI for code generation has transformed software development, but it has also introduced a critical transparency problem: the origins of AI-generated code are opaque to the developers who use it, the organizations that deploy it, and the compliance professionals responsible for ensuring its legal and quality standards. Existing mitigations flag problematic outputs after the fact without explaining why a model produced them or how future generation could be improved. We present a research vision, grounded in a U.S. NSF-funded research grant, that argues that the next generation of CodeGenAI tools must be built on a foundation of explainable provenance: automated, post-hoc traceability that links generated code back to the prompt components, training data instances, global data features, and internal model components that caused its generation. We grounded this vision in empirical evidence from studies of software developers, model users, and compliance/legal professionals, which show that provenance information is a practical necessity that current tools do not provide. We characterize the problem across four traceability dimensions, outline a research program combining large-scale empirical studies with post-hoc causal and interpretability techniques, and identify the key open challenges that the community must address to realize this vision.

↑