发表机构
University College London(伦敦大学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过医疗平台案例,探讨AI编写软件中人类控制的可靠性,发现测试、监控和审查智能体存在缺陷,提出将结果、证据、权限和决策统一于同一目标以保障人类控制。
AI 中文摘要
软件工程智能体能够使没有正式软件培训的人构建他们原本无法实现的系统,同时也能产生比专家能够有意义检查的更多的代码。在这两种情况下,详尽的代码审查作为人类控制的唯一基础并不可靠。我们报告了一个通过编码智能体构建并由没有正式软件工程培训的操作员管理的生产医疗保健平台的案例研究。随着时间的推移,其工作流程演变成一个以人类为主导的元智能体系统,其中一个智能体编写代码,其他智能体监督和审查代码,项目规则将经验教训向前传递。操作员发现用于监督系统的测试、监控器和审查智能体都可能出错。一些监控器测量的是代理指标而非结果,一些审计静默失败,缺失的检查从报告结果中消失,一次自动修复导致了运营中断。在这个案例中,人类控制依赖于将预期结果、用于判断结果的证据、智能体的权限和最终的人类决策都绑定到同一底层目标上。
英文摘要
Software-engineering agents can enable people without formal software training to build systems they could not otherwise implement and simultaneously can produce more code than even experts can meaningfully inspect. In both cases, exhaustive code review is not reliable as the sole basis for human control. We report a case study of a production healthcare platform built through coding agents and governed by an operator without formal software-engineering training. Over time, its workflow grew into a human-led meta-agent system where one agent wrote code, other agents supervised and reviewed it, and project rules carried lessons forward. The operator found that tests, monitors and reviewing agents used to supervise the system were fallible. Some monitors measured proxies rather than outcomes, some audits failed silently, missing checks disappeared from reported results and one automated repair caused operational disruption. In this case, human control depended on keeping the intended outcome, the evidence used to judge it, the agents' permissions and the final human decision were all tied to the same underlying objective.
CommentsAccepted to the NeurIPS 2026 Meta-Agents Workshop