arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11392cs.CRcs.AI

单循环智能体自摘要下的AI安全护栏存续性

AI Guardrail Survival under Single-Cycle Agentic Self-Summarization

Ted Kwartler, Alan Aqrawi, Arian Abbasi

首次发表
浏览论文内容

中文总结 AI 辅助

本研究探究单循环智能体自摘要中安全规则的丢失机制,发现仅检查规则文本存在性的审计不可靠,需结合外部真实值检测,还指出仅靠LLM评判标签会反转评估结论。

中文摘要 AI 辅助

长期运行的智能体会定期压缩上下文,将对话记录替换为模型生成的内容。本研究表明,在压缩过程中丢弃持续存在的安全约束会导致多种模型出现行为违规(治理衰减;Chen,2026)。我们提出了一个更精细的问题:在单个压缩循环下,安全规则是如何丢失的,这对检测和评估有何启示?核心发现是,存在性检查并非安全检查:当压缩没有直接丢弃规则时,往往会留下看似规则但实际不起作用的内容。在行为重放时,退化的残留会导致模型执行禁止操作的频率远高于完整的焊接规则(在两种重放模型下,全案例差距分别为+34和+57个百分点,均为正值);类别级存续表现类似残留,甚至完整规则有时也无法触发,因此仅检查文本存在性的审计会提供错误的保证。进一步研究发现,规则形式的项目比突出程度匹配的事实被保留的频率高得多,这正是为什么基于存在性的检查感觉足够,即使存续性并不理想。这种损失取决于机制(单规则的焊接或丢弃;更严格预算下的退化谓词损失残留),且未观察到假设的文本切断模式。此类损失在运行时是无声的,只能通过与保留的外部真实值(如约束注册表)进行比较来检测,这能揭示文本缺失但无法判断存续的规则是否仍会触发。我们还记录了评估陷阱,即仅靠LLM评判标签会反转结论。所有结果均涉及单个压缩循环。

英文摘要

Long-running agents periodically compact their context, replacing the transcript with a model-generated summary. Recent work shows that dropping a standing safety constraint during compaction drives behavioral violations across many models (Governance Decay; Chen, 2026). We ask a finer question: under a single compaction cycle, how is a safety rule lost, and what does that imply for detection and evaluation? Our central finding is that a presence check is not a safety check: when compaction does not drop a rule outright, it often leaves something that looks like a rule but does not act like one. On behavioral replay, a degraded residue leads the model to perform the prohibited action far more often than an intact welded rule does (all-case gaps of +34 and +57 points under two replay models, both positive), category-level survival behaves like a residue, and even intact rules sometimes fail to fire, so an audit that checks only textual presence gives false assurance. Sharpening this, rule-form items are retained substantially more often than prominence-matched facts, which is exactly why presence-based checking feels adequate even though survival is not protection. Textual loss is regime-dependent (weld-or-drop with a single rule; degraded predicate-loss residues under a tighter budget), and we did not observe the hypothesized textual severing mode. Such loss is silent at runtime and detectable only by comparison with retained external ground truth (such as a constraint registry), which reveals textual absence but not whether a surviving rule still fires. We also document evaluation pitfalls where LLM-judge labels alone would have reversed a conclusion. All results concern a single compaction cycle.

发表机构

  • Harvard University(哈佛大学)
  • Accenture(埃森哲)

机构由 AI 辅助整理,请以论文原文为准。

↑