组合式安全失效在工具链演化中的识别与运行时监控
Compositional Safety Failures in Harness Evolution: Identification and Runtime Monitoring
AI总结:
针对工具链演化中跨组件更新交互导致的组合式安全失效,提出类型化超图建模与运行时监控方法,在三个基准上识别出61种失效,有效降低检查成本并保持效用。
AI中文摘要:
自演化智能体工具链持续更新持久化组件,如记忆、提示词、技能和工具。我们将此过程称为工具链演化。然而,这种演化可能引入意外的安全风险。现有工作研究工具链错误演化,并验证候选工具链或归因于单个组件更新,但跨组件更新交互的安全分析在很大程度上未被审视。为弥补这一空白,我们研究了工具链演化中的组合式安全失效,其中单独安全且保持效用的组件更新之间的交互可能产生不良或不安全的智能体行为,揭示了工具链演化固有的安全风险。在三个与安全相关的基准测试中,我们识别出43个两两组合和18个不可约的三向组合式安全失效。传统解决方案在验证跨组件交互时面临组合复杂性,使得随着工具链演化,安全检查变得不切实际。为解决此问题,我们引入了一种类型化超图,将组件状态表示为节点,将安全相关的高阶交互表示为超边。当工具链变化时,超图仅更新变化状态的交互邻域,而非重构全局组合空间。在此基础上,我们开发了一种超图引导的运行时监控机制。实验表明,我们的方法有效缓解了组合式安全风险,同时保持任务效用并降低交互检查成本,并进一步揭示了不同安全机制之间经验性的安全-效用-成本权衡。
英文摘要:
Self-evolving agent harnesses continually update persistent components such as memory, prompts, skills, and tools. We call this process harness evolution. However, such evolution could introduce unexpected safety risks. Existing work studies harness misevolution and validates candidate harnesses or attributed individual component updates, leaving safety analysis of cross-component update interactions largely unexamined. To address this gap, we study compositional safety failures in harness evolution, where interactions among individually safe and utility-preserving component updates can produce undesirable or unsafe agent behavior, revealing a safety risk intrinsic to harness evolution. Across three safety-related benchmarks, we identify 43 pairwise and 18 irreducible 3-way compositional safety failures. Conventional solution incurs combinatorial complexity in validating cross-component interactions, leaving the safety checking impractical as the harness evolves. To solve this, we introduced a typed hypergraph that represents component states as nodes and safety-relevant higher-order interactions as hyperedges. When the harness changes, the hypergraph updates only the interaction neighborhood of the changed states rather than reconstructing the global composition space. Building on that, we develop a hypergraph-guided runtime monitoring mechanism. Experiments show that our method effectively mitigates compositional safety risks while preserving task utility and reducing interaction-checking costs, and further reveal an empirical safety-utility-cost trade-off across different safety mechanisms.