AI 中文总结
提出CodeSpec双可执行规范方法,通过配对子需求语义与仓库架构构建可靠功能链,在FeatureBench上优于基线,提升了长周期功能开发的设计与实现一致性及性能。
AI 中文摘要
基于大语言模型(LLM)的代码智能体通过与代码库及工具的迭代交互,已在仓库级软件开发中取得进展。然而,功能开发需要将新行为整合到现有架构中,形成连贯的跨组件功能链。现有智能体通常通过自由推理推导此类链,常产生功能链不完整的不可靠功能设计;且文本设计难以验证和执行,在长周期开发中难以维持设计与实现的一致性。本文提出CodeSpec,一种面向仓库级功能开发的双可执行规范方法:它通过将子需求语义与仓库架构配对的证据构建可靠功能链,再将其编译为互补的架构与行为规范,可在长交互中检查链的完整性与正确性,同时保持设计与实现的一致性。在针对现有仓库功能开发的FeatureBench上,CodeSpec在DeepSeek-V4-Pro下的通过率达70.7%、55.0%和49.9%,优于Claude Code等代表性基线;在仓库生成基准NL2Repo-Bench上的结果进一步证明了其通用性。
英文摘要
LLM-based code agents have advanced repository-level software development through iterative interaction with codebases and tools. However, feature development requires integrating new behaviors into existing architectures through coherent cross-component functional chains. Existing agents typically derive such chains through free-form reasoning, often producing unreliable feature designs with incomplete functional chains. Moreover, textual designs are difficult to verify and enforce, making it challenging to maintain design-implementation consistency throughout long-horizon development. We propose CodeSpec, a dual executable specification method for repository-level feature development. It builds reliable functional chains from evidence pairing sub-requirement semantics with repository architectures, then compiles them into complementary architecture and behavior specifications that check chain completeness and correctness while preserving design-implementation consistency over long interactions. On FeatureBench, which targets feature development in existing repositories, CodeSpec achieves 70.7%, 55.0%, and 49.9% pass rates under DeepSeek-V4-Pro, outperforming representative baselines such as Claude Code. Results on the repository generation benchmark NL2Repo-Bench further demonstrate its generalizability.