现实是最终验证者:论智能体软件工程中的两个关键差距
Reality Is the Final Verifier: On Two Key Gaps in Agentic Software Engineering
浏览论文内容
中文总结 AI 辅助
本文提出两差距框架,指出需求与模型差距是智能体软件工程失败主因,并引入保证-修订循环,以现实部署为最终验证,持续缩小差距。
中文摘要 AI 辅助
软件开发遵循一个实现-验证循环,在此循环中,开发者或智能体迭代地修改实现,直到评估器(如测试套件)接受它。评估器在部署环境的模型下,根据一组需求检查实现。然而,即使有形式化证明表明该实现在模型下满足需求,也无法保证部署后的行为是可接受的。需求仅近似于利益相关者的意图,而模型仅近似于真实的部署环境。我们将这两者统称为需求差距和模型差距——即两差距框架,该框架统一了智能体软件工程的主要失败模式:奖励黑客利用需求或模型中的遗漏,而幻觉则通过捏造需求或环境假设来扩大差距。由于在开放、变化的世界中,这两个差距通常无法被证明已闭合,目标从闭合它们转变为持续缩小它们。因此,我们提出一个保证-修订循环,当利益相关者拒绝由此产生的行为时,该循环利用部署证据来修订需求、模型或评估器。然后,我们将有保证的智能体开发视为一个在人类判断、智能体能力和计算资源之间的资源分配问题。两个主要瓶颈与两个差距相对应:人类判断对应需求差距,忠实且昂贵的评估对应模型差距。现实仍然是最终验证者:实际部署条件下可接受的行为是最终测试,而部署前评估只是其代理。
英文摘要
Software development follows an implementation-verification loop in which developers or agents iteratively revise an implementation until an evaluator, such as a test suite, accepts it. The evaluator checks the implementation against a set of requirements under a model of the deployment environment. Yet even a formal proof that the implementation satisfies the requirements under the model cannot guarantee acceptable behavior after deployment. Requirements only approximate stakeholder intent, and the model only approximates the real deployment environment. We call these together - requirement gap and model gap - the two-gap framework, which unifies the main failure modes of agentic software engineer-ing: reward hacking exploits omissions in the requirements or model, while hallucination widens the gaps by fabricating requirements or environment assumptions. Because neither gap can generally be certified closed in an open, changing world, the goal shifts from closing them to continuously narrowing them. We therefore propose an assurance-revision loop that uses deployment evidence to revise the requirements, model, or evaluator when stakeholders reject the resulting behavior. We then cast assured agentic development as a resource-allocation problem over human judgment, agent capability, and compute. The two principal bottlenecks mirror the two gaps: human judgment for the requirement gap and faithful, costly evaluation for the model gap. Reality remains the final verifier: acceptable behavior under actual deployment conditions is the ultimate test, while predeployment evaluations remain proxies for it.
发表机构
- UC Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。