arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05531cs.CLcs.AIcs.CYcs.MA

何时智能体治理有帮助

When Agent Governance Helps

  • Purdue University(普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

Michael Ray Johnson, Linda Naimi

AI总结:

本研究提出GAMPO框架并实验验证,发现治理收益受模型备用容量限制,基于案例的规范优于统一流程,能显著提升任务成功率。

AI中文摘要:

没有任何规范说明应如何设计和评估一个受治理的自主目标智能体组织,其中智能体在护栏内追求自生成目标。我们分两部分回答这个问题。首先,我们通过对321个来源的基于文档的定性证据综合,综合出受治理的自主多智能体产品组织(GAMPO)框架,将能动性、敏捷、平台和治理理论整合为一个可运行的规范。其次,我们在CHI-Bench(一个长视野医疗保健基准)上,跨开放模型和前沿模型,探究了GAMPO的提示层实例化。结果是一个边界条件:治理收益受模型备用容量的门控,并且是领域和模型特定的。在容量受限的开放模型上,完整流程没有产生可靠收益,而一句“验证你的写入”将任务成功率翻倍(pass@1从2/20提高到4/20)。在前沿模型上,相同的脚手架将预先授权从24%提升到40%,但在另一个模型上净收益为零,这一差距可追溯到一种稳定的推荐覆盖倾向。第二个结果细化了第一个结果:用基于案例自身政策和已发布标准(而非隐藏键)的、对答案不知情的、每任务完成定义替换通用流程,在最佳五次自一致性下将预先授权提高到84%(单次尝试为68%,由保留委员会确认),并将利用管理提高到44%,而护理管理则遇到主观内容质量壁垒。贡献是一个命名、可审计的框架和基于能力门控的证据,表明治理应适应备用容量,并且在前沿,基于案例的规范优于统一流程。研究结果是探索性的:部分实例化、每单元样本量小(n=5-25)和单次试验。

英文摘要:

No specification says how a governed autotelic AI agent organization, where agents pursue self-generated goals inside guardrails, should be designed and evaluated. We answer in two parts. First, we synthesize the Governed Autotelic Multi-Agent Product Organization (GAMPO) framework from a document-based qualitative evidence synthesis of 321 sources, integrating agency, agile, platform, and governance theory into a runnable specification. Second, we probe a prompt-layer instantiation of GAMPO on CHI-Bench, a long-horizon healthcare benchmark, across open and frontier models. The result is a boundary condition: governance benefit is gated by a model's spare capacity and is domain- and model-specific. On capacity-constrained open models the full procedure yields no reliable benefit, whereas a single "verify your writes" sentence doubles task success (pass@1 2/20 to 4/20). At the frontier the same scaffold lifts prior-authorization 24% to 40% but nets zero on another model, a gap traced to a stable recommendation-override disposition. A second result refines the first: replacing the generic procedure with an answer-blind, per-task definition-of-done, keyed only to the case's own policy and published standards, never the hidden key, raises prior-authorization to 84% under best-of-five self-consistency (68% single-attempt, confirmed by a held-out board) and utilization-management to 44%, while care-management meets a subjective content-quality wall. The contribution is a named, auditable framework and capability-gated evidence that governance should be sized to spare capacity, and that at the frontier a case-grounded specification beats a uniform procedure. Findings are exploratory: partial instantiation, small per-cell samples (n = 5-25), and single trials.

补充信息

↑