arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大规模仓库工程:基于智能体原生可复用代码原语

Large-scale Repository Engineering via Agent-Native Reusable Code Primitives

Haibo Jin, Peng Kuang, Xucheng Yu, Jerry Wang, Dehao Wu, Haohan Wang

arXiv 2610.09079首次发表:更新:

发表机构

University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出代码原语和LEGO框架,通过可复用可执行组件库CodeFace实现仓库级代码构建,在LEGO-REPO基准上显著提升13个骨干模型的交付分数,最高提升61.4%。

AI 中文摘要

配备开发环境的大型语言模型已将代码生成推向仓库级构建,但由于交互模块、接口、配置、测试和依赖必须协同工作,构建完整仓库仍然困难。我们引入代码原语(Code Primitives),即具有接口契约、依赖闭包、验证测试和溯源信息的智能体原生可复用可执行组件。每个原语使用常驻LLM评估相关性并调整其实现、接口和依赖以适应目标仓库,我们将1,424个经过验证的原语组织在CodeFace中,这是一个用于仓库构建的可搜索库。我们提出LEGO(通过智能体原生可复用代码原语进行大规模仓库工程),它激活任务相关原语,将其调整后的实现与任务特定代码集成,同时解决跨组件约束,并根据执行的测试修订结果。为端到端衡量构建效果,我们构建了LEGO-REPO基准,包含522个可执行重建任务,涵盖七个软件领域、22个能力轨道和五个难度级别,评分依据原生测试套件,介于空包下限和原始源码上限之间。13个评估骨干中最强者达到0.318的交付分数,在41.0%的任务上得分为零;LEGO平均提升所有13个骨干0.1474,并将GPT-5.6-terra从0.3180提高到0.5134(+61.4%)。在受控比较中,调整后的原语优于作为上下文提供或未经修改的检索代码。该效果在独立仓库代理、三个外部基准以及重新挖掘的CodeFace上持续存在;用于调整和诊断的GPT-OSS-20B在成本降低24.0%的情况下保留了95.1%的同质分数。

英文摘要

Large language models equipped with development environments have moved code generation toward repository-scale construction, yet building complete repositories remains difficult because interacting modules, interfaces, configurations, tests, and dependencies must work together. We introduce Code Primitives, agent-native reusable executable components with interface contracts, dependency closures, validation tests, and provenance. Each primitive uses a resident LLM to assess relevance and adapt its implementation, interfaces, and dependencies to the target repository, and we organize 1,424 validated primitives in CodeFace, a searchable library for repository construction. We introduce LEGO (Large-scale repository Engineering via aGent-native reusable cOde primitives), which activates task-relevant primitives, integrates their adapted implementations with task-specific code while resolving cross-component constraints, and revises the result against executed tests. To measure construction end to end, we build LEGO-REPO, a benchmark of 522 executable reconstruction tasks spanning seven software domains, 22 capability tracks, and five difficulty levels, scored against native test suites between an empty-package floor and original-source ceiling. The strongest of 13 evaluated backbones reaches a delivery score of 0.318 and scores zero on 41.0% of tasks; LEGO improves all 13 by 0.1474 on average and raises GPT-5.6-terra from 0.3180 to 0.5134 (+61.4%). In controlled comparisons, adapted primitives outperform retrieved code supplied as context or vendored unchanged. The effect persists against independent repository agents, across three external benchmarks, and with a disjointly re-mined CodeFace; GPT-OSS-20B for adaptation and diagnosis retains 95.1% of the homogeneous score at 24.0% lower cost.

Comments44 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑