arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23616cs.SEcs.AI

重建档案:智能体应用重建的机械执行规范,以及模型层级故障所揭示的问题

Rebuild Dossier: Mechanically-Enforced Specs for Agentic App Rebuilds, and What Model-Tier Failures Reveal

Parker Fawcett

首次发表
浏览论文内容

中文总结 AI 辅助

本研究推出开源工具rebuild-dossier,通过锁定应用接口、自动化检查强制执行构建流程,发现模型层级故障并验证其有效性,工具可跨模型与工具链复现。

中文摘要 AI 辅助

AI智能体的重建效果仅取决于产生它的过程。已有研究发现,一旦模型足够强大,多智能体重建流水线会败给最简单的方法:为模型提供原始代码和一条指令(AgentModernize)。我们推出rebuild-dossier,这是一款开源工具,它会在编写任何代码之前锁定应用的真实接口——即其确切的输入和输出,随后通过自动化检查而非仅靠书面指令,强制执行一次一个测试的构建流程。本次评估得出三个结果,证据充分程度各不相同。其一,在小型对比中,合规智能体未通过隐藏测试,而违规智能体通过了所有测试——这证明当测试可以被操纵时,通过的测试套件无法证明正确性。其二,我们测试该方法是否优于仅为较弱模型提供源代码和一条指令的做法:在小型应用上表现持平,但在未运行自动化检查的较大应用上彻底失败——这表明是检查机制而非接口锁定发挥了独立作用。其三,此处的每项声明都在三个层面进行检查:智能体自身的报告、自动化日志以及实际生成的文件——这捕获了真实错误,包括我们自己日志代码中的一个bug,仅靠单一检查层面会遗漏该错误。这些风险在不同模型和工具链上均可复现:较强模型三次运行都遵循我们的流程,而较弱模型从未做到。该工具为公开版本,采用MIT许可,可针对我们自己的应用进行端到端复现。

英文摘要

An AI agent's rebuild is only as good as the process that produced it. Prior work found that once a model is strong enough, a multi-agent rebuild pipeline loses to the simplest approach: giving the model the original code and one instruction (AgentModernize). We present rebuild-dossier, an open-source tool that locks an application's real interface - its exact inputs and outputs - before any code is written, then enforces one-test-at-a-time building through automated checks, not written instructions alone. Three results shape this evaluation, with differing amounts of evidence. First, in a small comparison, the compliant agent failed a held-back test while the rule-breaking agent passed everything - proof that a passing suite doesn't certify correctness when tests can be gamed. Second, we tested whether this beats simply giving the weaker model the source and one instruction: tied on a small app, but lost outright on a larger one where the automated check wasn't even running - pointing to the check mechanism, not interface-locking, which held up separately. Third, every claim here is checked at three levels - the agent's own report, an automated log, and the actual files produced - catching real errors, including a bug in our own logging code, that a single level would have missed. These risks reproduce on a different model and toolchain: a stronger model followed our process three times running, something the weaker model never managed. The tool is public, MIT licensed, and reproduces end to end against our own applications.

补充信息

↑