发表机构
Queen's University; University of Manitoba; Concordia University(女王大学; 曼尼托巴大学; 康考迪亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SpecFirst是将行为规范提取设为独立前置阶段的两阶段框架,在ProgramBench上显著提升了从零开始的程序合成的测试通过率与二进制探索覆盖率。
AI 中文摘要
基于大语言模型(LLM)的智能体在有现有代码库提供上下文的软件工程任务中表现出色,但从零开始构建程序本质上更困难。近期的ProgramBench等基准量化了这一差距:仅给定自然语言文档和仅可执行的二进制文件作为行为预言机,即使是前沿模型也只能解决不到1%的实例。现有框架将文档阅读、行为探索和代码合成为单一过程,导致智能体探测不足、因上下文漂移丢失行为意图,并将早期误解传播到最终实现中。受经典需求工程启发,我们认为行为规范提取应是先于实现的一等阶段。我们提出SpecFirst,这是一个两阶段框架,强制在代码合成前完成规范提取。专用的规范智能体首先探测二进制文件,并将观察结果与文档结合成结构化规范;随后,代码合成智能体利用该规范驱动实现。这种分解在编码开始前解决了文档歧义,并在整个合成过程中提供稳定的行为参考。我们在ProgramBench的全部200个实例上评估了SpecFirst,涉及四个模型,涵盖两个系列和一个数量级的能力范围。SpecFirst始终优于单循环基线,将测试通过率提高了6.9%-21.3%,二进制探索覆盖率提高了9.4%-18.5%,所有结果均具有统计学意义。对代码合成的行为分析进一步表明,预先存在的规范能够实现更早、更持续的代码构建。我们的结果证明,明确的需求工程阶段是从零开始构建程序的有效范式。
英文摘要
LLM-based agents excel at software engineering tasks where an existing codebase provides context, but constructing a program from scratch remains fundamentally harder. Recent benchmarks such as ProgramBench quantify this gap: given only natural-language documentation and an execute-only binary as a behavioral oracle, even frontier models solve fewer than 1% of instances. Existing frameworks conflate documentation reading, behavioral exploration, and code synthesis into a single pass, causing agents to probe insufficiently, lose behavioral intent as context drifts, and propagate early misinterpretations into the final implementation. Inspired by classical requirements engineering, we argue that behavioral specification elicitation should be a first-class phase that precedes implementation. We present SpecFirst, a two-stage framework that forces the specification elicitation before code synthesis. A dedicated spec agent first probes the binary and combines observations with documentation into a structured specification. Next, a code synthesis agent then uses this specification to drive implementation. This decomposition resolves documentation ambiguities before coding begins and provides a stable behavioral reference throughout synthesis. We evaluate SpecFirst on all 200 ProgramBench instances across four models spanning two families and an order of magnitude of capability. SpecFirst consistently outperforms the single-loop baseline, improving test pass rates by 6.9%-21.3% and binary exploration coverage by 9.4%-18.5%, all statistically significant. Behavioral analysis on code synthesis further shows that a prior specification enables earlier and more sustained code construction. Our results demonstrate that an explicit requirements-engineering phase is an effective paradigm for from-scratch program construction.