AI 中文总结
研究使用工具的智能体认证运行时安全,区分三个问题:确定性门与安全策略、外部法则下的边界及证书、闭环边界识别;介绍了相关方法,通过多种实验针对差异进行研究。
AI 中文摘要
运行时护栏在不可逆转的工具调用之前起作用,但其保证取决于可表示的策略状态、法官观察到的内容以及干预是否会改变未来行为。我们区分了三个问题。首先,相对于固定的预言机谓词,确定性门精确地强制执行其寄存器模型识别出良好前缀的非空安全策略;对于两个可递减计数器,策略非平凡性是不可判定的,但对于可分离单调片段则在PSPACE中。其次,在固定的外部法则下,奈曼 - 皮尔逊给出了确切的错误阻止/遗漏边界,共形校准给出了有限样本边际证书,可能通过全块方式。第三,一旦阻止改变未来提议,静态分数和无门控轨迹不一定能识别闭环边界;指定的有限受控模型反而会产生占用程序。有界表示攻击增加了鲁棒性余量,所以仅良性校准无法转移。实验通过静态诊断、受控模型枚举、表示重写和配对闭环重新运行来针对这些差异。
英文摘要
Runtime guardrails act before irreversible tool calls, but their guarantees depend on what policy state is representable, what a judge observes, and whether intervention changes future behavior. We separate three questions. First, relative to fixed oracle predicates, a deterministic gate enforces exactly the nonempty safety policies whose good prefixes its register model recognizes; policy nontriviality is undecidable with two decrementable counters but in PSPACE for a separable monotone fragment. Second, under a fixed exogenous law, Neyman-Pearson gives the exact false-block/miss frontier and conformal calibration gives a finite-sample marginal certificate, possibly via block-all. Third, once blocking changes future proposals, static scores and ungated trajectories need not identify the closed-loop frontier; a specified finite controlled model instead yields an occupancy program. Bounded representation attacks add a robustness margin, so benign calibration alone does not transfer. Experiments target these distinctions through static diagnostics, controlled-model enumeration, representation rewrites, and paired closed-loop reruns.
Comments26 pages, 8 figures. Extended version with complete proofs and additional experiments