arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12216cs.RO

护栏元代理循环:策略固定、预算边界与崩溃恢复的压力测试

Guardrailed Meta-Agent Loops: Stress-Testing Policy Pinning, Budget Bounds, and Crash Recovery

Qinzhen Ma, Jialin Wu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出GuardrailLoop仿真测试平台,联合测试策略固定、计算核算和崩溃恢复三个契约,通过50种子实验和240次崩溃注入证明结果恢复不等于恰好一次执行。

中文摘要 AI 辅助

自我改进的代理工作流在同一个控制器既能改变其行为又能改变评判该行为的条件时,会产生审计问题。我们提出了GuardrailLoop,一个基于仿真的测试平台,使三个操作契约可联合测试:人类定义策略的保留、每个记录执行前缀的计算核算,以及崩溃后恢复指定科学状态。哈希固定的策略固定目标、范围、评估身份、预算和发布条件;机器引导的进化被限制在代码拥有的特征目录和有界旋钮内。贡献在于一个可执行边界和一个评估协议,该协议区分了有用的适应、状态恢复和重复执行。在一项配对的50种子2x2研究中,回合阶段增长将目标达成率提高了+1.00,并将达到目标的受限平均计算量减少了-56.97模拟GPU小时(95%配对自助法区间[-58.91,-54.70]);空闲增长对效用的测量影响为零。在240次枚举的崩溃注入中,所有运行都恢复了定义的结果,但只有210次保留了规范化轨迹:30次提交前崩溃重复了一次规划器调用。资源漂移、终止开关、完整性和输出护栏矩阵均满足其规定的检查。这些发现表明,成功的结果恢复不足以证明恰好一次执行。它们在一个校准的确定性测试平台内建立了符合性,而非通用安全性或现实世界的自我改进。

英文摘要

Self-improving agent workflows create an audit problem when the same controller can change both its behavior and the conditions under which that behavior is judged. We present GuardrailLoop, a simulation-based testbed that makes three operational contracts jointly testable: preservation of human-defined policy, compute accounting at every recorded execution prefix, and recovery of a specified scientific state after crashes. A hash-pinned policy fixes goals, scope, evaluation identity, budget, and release conditions; machine-directed evolution is restricted to a code-owned feature catalog and bounded knobs. The contribution is an executable boundary and an evaluation protocol that separates useful adaptation, state recovery, and repeated execution. In a paired 50-seed 2 x 2 study, round-stage growth changes target attainment by +1.00 and restricted mean compute to target by -56.97 simulated GPU-hours (95% paired-bootstrap interval [-58.91,-54.70]); idle growth has zero measured utility effect. Across 240 enumerated crash injections, all runs recover the defined outcome, but only 210 preserve the normalized trace: 30 pre-commit crashes repeat a planner call. Resource-drift, kill-switch, integrity, and output-guard matrices satisfy their specified checks. These findings show why successful outcome recovery is insufficient evidence of exactly-once execution. They establish conformance within one calibrated deterministic testbed, rather than general safety or real-world self-improvement.

发表机构

  • Rice University(莱斯大学)
  • University of California, San Diego(加利福尼亚大学圣迭戈分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑