arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10961cs.SEcs.AI

跨提供商审查作为编码智能体的运行时契约:一项受控试点与故障注入研究

Cross-Provider Review as a Runtime Contract for Coding Agents: A Controlled Pilot and Fault-Injection Study

Bowen Xu, Boyu Chen

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出跨提供商审查契约作为编码智能体的运行时契约,通过受控试点、故障注入等测试发现了多个缺陷,为编码智能体的运行时可靠性提供了验证方案。

中文摘要 AI 辅助

编码智能体越来越多地共享工作站,同时依赖不同的提供商和订阅额度。第二个智能体可以检查已完成的答案,但调用会消耗另一部分额度,且可能无法提供实质性发现。我们描述了一种咨询型跨提供商审查契约:包含独立资源池、受限执行、受限审查员能力、完整输入交付、可用语义输出、明确失败状态以及每次尝试的持久证据。在一项受控的、由智能体生成的20次配对开发轮次试点中,8次出现了实质性的审查员发现(95%精确区间为19.1-63.9%)。对两个审查员后端进行的边界条件扫描重现了先前发现的部分输入下的假成功:四个截断级别在历史上通过,修复后失败。该扫描还发现并修复了进程回收期间的取消问题。在真实的CLI探测中,Claude没有写入工具;Codex在5次只读试验中均尝试写入,每次写入工具均失败,且无临时仓库发生变更。这些测试覆盖了指定路径和版本,而非现场可靠性。一项预注册的仅元数据审查分配的影子研究积累了25条正式观测结果,之后一次精确运行时回归发现了第三个缺陷:审查员以非零退出码退出且带有格式良好的裁决被计为完成。每次尝试未记录退出状态,因此无法追溯解决暴露问题。25条正式记录和2条待处理记录仍作为审计队列;测量有效队列已从零重启,且收集工作已开始。未报告任何门控结果。

英文摘要

Coding agents increasingly share a workstation while drawing on separate providers and subscription allowances. A second agent can inspect a completed answer, but the call spends another pool and may provide no substantive finding. We describe an advisory cross-provider review contract: distinct resource pools, bounded execution, restricted reviewer capabilities, complete input delivery, usable semantic output, explicit failure states and durable per-attempt evidence. In a controlled, agent-authored pilot of 20 paired development turns, eight had a material reviewer finding (95% exact interval 19.1-63.9%). A boundary-condition scan across both reviewer backends reproduced a previously discovered false success on partial input: four truncation levels passed historically and failed after repair. The scan also found and repaired cancellation during process reaping. In real CLI probes, Claude had no writing tools; Codex attempted writes in five of five read-only trials, each write tool failed, and no disposable repository changed. These tests cover specified paths and versions, not field reliability. A preregistered shadow study of metadata-only review allocation accrued 25 formal observations before an exact-runtime regression found a third defect: a reviewer exiting nonzero with a well-formed verdict was counted as complete. Exit status was not recorded per attempt, so exposure cannot be resolved retrospectively. The 25 formal and two pending records remain an audit cohort; the measurement-valid cohort restarted at zero and collection has begun. No gate result is reported.

发表机构

  • Stanford University(斯坦福大学)
  • University College London(伦敦大学学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑