arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ClawSentry:一种用于保护自主大语言模型智能体的渐进式多层安全监控器

ClawSentry: A Progressive Multi-Tier Security Monitor for Safeguarding Autonomous LLM Agents

Kai Wang, Zeming Wei, BiaoJie Zeng, Chang Jin, An Wang, Xiaokun Luan, Zhixiao Lin, Jingjing Qu, Xia Hu, Xingcheng Xu

arXiv 2608.21101首次发表:更新:

发表机构

Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ClawSentry是一种开源、与框架无关的智能体运行时安全监控网关,通过渐进式多层机制防护自主LLM智能体,可大幅降低攻击成功率,同时保持较高的任务安全率。

AI 中文摘要

随着大语言模型(LLM)智能体从对话转向执行代码、读取本地文件以及协调外部工具,被恶意第三方技能劫持的单个智能体可能会导致数据外泄、权限提升或级联破坏。我们认为智能体风险是渐进式的:它可在智能体控制循环的四个位点(技能准入、调用时意图、执行时效果、行动后后果)侵入,而被拒绝的危险目标可能以不同的表面形式、工具或回合重新出现;现有的安全措施通常仅针对一个生命周期边界或一次调用。基于该威胁模型,我们提出ClawSentry,这是一种开源、与框架无关的智能体运行时安全监控网关。在技能包执行前,首次使用技能包审查(FSPR)会基于确定性证据阈值对其进行审计,将未解决的案例升级为有界只读智能体审查(位点A);在运行时,一个三层渐进式决策引擎——确定性L1层、基于规则的L2语义审查器、只读L3证据搜寻智能体——仅对残留歧义进行上下文审查,同时会话级反绕过机制识别工具切换和改写后的重试(位点B-C);行动后路径将高严重性证据非追溯性地输入后续审查(位点D)。智能体约束协议(AHP)抽象可在不修改智能体内部结构的情况下,对Codex、Claude Code、Kimi CLI和Gemini CLI应用统一策略。在SkillInject数据集上,使用Codex/GPT-5.4时,上下文攻击成功率(ASR)从39.55%降至2.61%,上下文任务安全率(TSR)仅从83.78%降至83.05%;在完整SkillsSafety基准的5个工作智能体上,未受保护时ASR为33.5%-49.7%,ClawSentry将其限制在9.09%-15.03%,干净技能的总TSR保持98.7%。

英文摘要

As large language model (LLM) agents move from conversation to executing code, reading local files, and orchestrating external tools, a single agent hijacked by a malicious third-party skill can cause data exfiltration, privilege escalation, or cascading compromise. We argue that agentic risk is progressive: it can enter at four loci of the agent control loop--skill admission, invocation-time intent, execution-time effect, and post-action consequence--while a denied dangerous objective can reappear across surface forms, tools, or turns; existing safeguards are typically local to one lifecycle boundary or one call. Guided by this threat model, we present ClawSentry, an open-source, framework-agnostic security supervision gateway for agent runtimes. Before a skill package is ever executed, First-use Skill Package Review (FSPR) audits it under a deterministic evidence floor, escalating unresolved cases to bounded read-only agentic review (locus A). At runtime, a three-tier progressive decision engine--a deterministic L1 layer, a rule-anchored L2 semantic reviewer, and a read-only L3 evidence-seeking agent--spends contextual review only on the residual ambiguity, while a session-level anti-bypass mechanism recognizes tool-switching and rephrased retries (loci B--C); a post-action path feeds high-severity evidence non-retroactively into later review (locus D). An Agent Harness Protocol (AHP) abstraction applies one policy across Codex, Claude Code, Kimi CLI, and Gemini CLI without modifying agent internals. On SkillInject with Codex/GPT-5.4, contextual ASR falls from 39.55% to 2.61% while contextual TSR moves only from 83.78% to 83.05%. Across five Work Agents on the full SkillsSafety benchmark, ClawSentry confines ASR to 9.09--15.03% from 33.5--49.7% unprotected, and aggregate TSR on clean skills remains 98.7%.

Comments35 pages, 14 figures. Code: https://github.com/Elroyper/ClawSentry

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑