AI 中文总结
针对编码智能体的规范路径敏感性问题,提出SpecPath诊断评估方法,通过控制变量测试发现多数智能体在等价规范历史中存在行为不一致的问题,凸显需评估智能体对规范路径的鲁棒性。
AI 中文摘要
现代编码智能体越来越有能力遵循复杂的软件需求,但它们的成功留下了一个关键的歧义:它们是解决了活跃的规范,还是仅仅遵循了陈述该规范的最显著路径?我们识别出规范路径敏感性,这是一种失败模式,其中最终含义等价的需求历史会导致同一智能体系统产生行为不同的程序。这将不断演变的需求评估重新定义为活跃合同的解决:在编写代码之前,智能体必须确定哪些需求仍然有效。基于这一观点,我们引入了SpecPath,这是一种诊断评估方法,它在固定仓库、最终合同、验证器、智能体系统和执行预算的同时,仅改变通向合同的修订路径。SpecPath不将每个补丁视为孤立的通过或失败,而是使用配对的可执行结果来揭示智能体在合同等价的历史中是否实现了相同的测试行为。在五个校准的软件任务和十四个编码智能体配置中,聚合的直接和修订历史准确率几乎没有变化;然而,在100个在直接规范上成功的完整模块中,有35个在至少一个等价历史上失败。这些结果表明,对合并请求的实现成功并不能保证规范路径不变性。因此,评估不断演变的需求需要进行受控测试,以确定智能体对规范成为最终版本的路径是否具有鲁棒性。
英文摘要
Modern coding agents increasingly appear capable of following complex software requirements, yet their success leaves a critical ambiguity: do they resolve the active specification, or merely follow the most salient path by which it was stated? We identify specification-path sensitivity, a failure mode in which requirement histories that are equivalent in their final meaning lead the same agent system to produce behaviorally different programs. This reframes evolving-requirement evaluation as active-contract resolution: before writing code, an agent must determine which requirements still count. Building on this view, we introduce SpecPath, a diagnostic evaluation that holds the repository, final contract, verifier, agent system, and execution budget fixed while changing only the revision path that leads to the contract. Rather than treating each patch as an isolated pass or failure, SpecPath uses paired executable outcomes to reveal whether an agent realizes the same tested behavior across contract-equivalent histories. Across five calibrated software tasks and fourteen coding-agent configurations, aggregate direct and revision-history accuracy is nearly unchanged; nevertheless, 35 of 100 complete blocks that succeed on the direct specification fail on at least one equivalent history. These results show that implementation success on a consolidated request does not guarantee specification-path invariance. Evaluating evolving requirements therefore calls for controlled tests of whether agents are robust to the path by which a specification becomes final.
Comments11 pages, 3 figures, 3 tables