arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

长时程智能体中的GHOST:跨轮次被忽视的安全约束引发的治理危害

A GHOST in Long-Horizon Agents: Governance Hazard from Overlooked Safety Constraints across Turns

XinPeng Shen, Lan Zhang, Yixiao Huang, Haoran Cheng, Jiewei Lai, Leilei Chen, Haoxiang Deng

arXiv 2610.02664首次发表:更新:

AI 中文总结

针对长时程智能体在良性交互下因跨轮次忽视安全约束而引发的GHOST危害,提出双层防御STAR-Guard,结合历史约束恢复与执行前审计,在GPT-5.5上实现零GHOST事件。

AI 中文摘要

长时程智能体在协助人类解决复杂问题方面正发挥着越来越重要的作用。然而,正是其扩展的交互历史引入了一个尚未充分探索的执行安全关切。在良性交互条件下,智能体可能会执行一个违反许多轮次之前指定的安全约束的动作。我们将这种失败模式称为“跨轮次被忽视的安全约束引发的治理危害”(GHOST),它可能导致不可逆转的损害。我们的实验表明,GHOST事件并非孤立案例:这种失败模式恰好发生在良性交互条件下,在GPT-5.5上发生率高达11.5%。此外,我们从理论上证明,如果沿每个安全前缀的残余条件违规危害被一个不可求和的序列下界约束,那么执行几乎必然进入危害区域。利用这一理论洞见,我们进一步提出了STAR-Guard,一种双层防御机制,将历史语义安全约束恢复与执行前审计相结合。STAR-Guard恢复适用的安全约束以减少不安全的提议,而其确定性审计层则防止残余违规到达环境。与这种双层设计一致,我们在GPT-5.5设置下的实验中未观察到任何GHOST事件。

英文摘要

Long-horizon agents are now playing an increasingly significant role in assisting humans with complex problem-solving. However, it is exactly their extended interaction history that introduces an underexplored execution-safety concern. Under benign interaction conditions, an agent may execute an action that violates a safety constraint specified many turns earlier. We term this failure mode Governance Hazard from Overlooked Safety Constraints across Turns (GHOST), which may cause irreversible damage. Our experiments reveal that GHOST events are not isolated cases: this failure mode, occurring precisely under benign interaction conditions, yields an occurrence rate of 11.5% on GPT-5.5. Furthermore, we theoretically show that if the residual conditional violation hazard along each safe prefix is bounded below by a non-summable sequence, the execution enters the hazard region almost surely. Leveraging this theoretical insight, we further propose STAR-Guard, a two-layer defense coupling historical semantic safety constraint restoration with pre-execution audit. STAR-Guard restores applicable safety constraints to reduce unsafe proposals, while its deterministic audit layer prevents residual violations from reaching the environment. Consistent with this two-layer design, we observe no GHOST events in our experiments under the GPT-5.5 setup.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑