arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当管理者成为干扰源:受管工具使用智能体中的控制生成干扰与成本感知退避

When the Governor Becomes the Disturbance: Control-Generated Disturbance and Cost-Aware Backoff in Governed Tool-Using Agents

Veronique Ziegler

arXiv 2610.09037首次发表:更新:

AI 中文总结

本研究在受控文件恢复环境中发现,不考虑成本的监管者会因干预强度增加而引发工具故障并阻塞任务,而基于已知诱发事件移动平均的自适应退避规则可降低干预概率、减少阻塞并提升完成率,实验验证了其有效性。

AI 中文摘要

监督管理者可能干扰其所监管的工具使用智能体。我们在一个受控的文件恢复环境中研究这种可能性,在该环境中,监管强度的增加会触发实验性强加的工具故障。一个不考虑成本的监管者可能将这些故障转变为持续阻塞,从而阻止任务完成。我们将该监管者与一种退避规则进行比较,该规则利用已知诱发事件的移动平均值来降低干预概率。在一个手工编码的随机策略智能体上,故障模式在结果替换和执行损坏的工具参数两种情况下均出现。对于持续策略,自适应退避相比干预频率大致匹配的固定弱监管者,提高了任务完成率。一个包含6个任务、共576个回合的Gemini 2.5 Flash实验也显示,在退避机制下阻塞减少且完成率提高;在测试的设置中,中等强度的退避取得了最高的总体成功次数。这些结果揭示了干预成本与持续动作阻塞之间的相互作用,并提供了一种可能的缓解措施。成本机制是强加的,其诱发事件对退避规则直接可观察;该机制在受控环境之外的适用性仍是一个实证问题。

英文摘要

Supervisory governors can interfere with the tool-using agents they regulate. We study this possibility in a controlled file-recovery environment where increases in regulatory intensity trigger experimentally imposed tool failures. A cost-blind governor can turn these failures into persistent blocking that prevents task completion. We compare this governor with a backoff rule that reduces intervention probability using a moving average of known induced events. On a hand-coded stochastic-policy agent, the failure pattern appears under both result replacement and execution of corrupted tool arguments. For the persistent policy, adaptive backoff improves completion relative to a fixed weak governor with approximately matched intervention frequency. A Gemini 2.5 Flash experiment comprising 576 episodes across 6 tasks also shows reduced blocking and improved completion under backoff; among the tested settings, intermediate backoff strength achieves the highest observed aggregate success. These results identify an interaction between intervention cost and persistent action blocking, together with a possible mitigation. The cost mechanisms are imposed and their induced events are directly observable to the backoff rule; applicability beyond this controlled environment remains an empirical question.

CommentsIncludes ancillary experimental data and a Python script for reproducing numerical summaries. Code: https://github.com/CriticalAttentionSystems/CASAgentGovernor

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑