arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当“不要”并非“拒绝”:CLAUDE.md中的安全规则与内置控制

When "Do Not" Is Not Deny: Security Rules in CLAUDE.md vs Built-In Controls

Ting Yan

arXiv 2608.23550首次发表:更新:

AI 中文总结

该研究对比CLAUDE.md的自然语言安全规则与Claude Code内置控制的匹配度,发现仅4%-16%的规则有匹配控制,揭示开发者无法获知规则是否被强制执行的安全问题。

AI 中文摘要

在CLAUDE.md中,“不要”是模型会解释的自然语言指令;Claude Code的拒绝(deny)是一种内置控制,会在智能体执行操作前将其阻止。两者都可表达相同的安全目标,但控制智能体的方式不同。我们在481个公开的CLAUDE.md文件中测量这种差距,让一个大语言模型(LLM)将提取的候选规则与Claude Code的文档化控制进行匹配,两名安全从业者在不查看模型答案及彼此标签的情况下独立检查样本。根据控制与书面规则的匹配紧密程度,仅约4%-16%的已检索安全规则有匹配的内置控制;在最严格标准下,估计值为4.4%(95%置信区间:2.6%-6.7%),两名标注者对哪些规则存在匹配的一致性很高。对完整文件的人工审查发现,我们的提取方法捕获了66.3%的符合条件的安全规则,因此报告的比率适用于其捕获的规则。这是一个可用的安全问题:CLAUDE.md是一个仅写通道,开发者编写安全规则却无法获得关于控制是否会强制执行的反馈;相同的纯文本形式隐藏了两类规则:可由权限规则、模式(mode)或沙箱(sandbox)强制执行的规则,以及留给模型解释的规则。

英文摘要

In CLAUDE.md, "do not" is a natural-language instruction that the model interprets. Claude Code's deny is a built-in control that blocks an action before the agent can take it. Both can express the same security goal, but they control the agent in different ways. We measure this gap in 481 public CLAUDE.md files. An LLM matched the extracted candidate rules against Claude Code's documented controls, and two security practitioners independently checked a sample without seeing the model's answers or each other's labels. Depending on how closely a control had to match the written rule, only about 4-16% of the retrieved security rules had a matching built-in control. Under the strictest standard the estimate was 4.4% (95% CI: 2.6-6.7%), and the two annotators agreed closely on which rules had a match. A manual review of complete files found that our extraction method captured 66.3% of eligible security rules; the reported rates therefore apply to the rules it captured. This is a usable security problem: CLAUDE.md is a write-only channel. A developer writes a security rule but gets no feedback on whether a control will enforce it. The same plain-text form hides two kinds of rule: those a permission rule, mode, or sandbox can enforce, and those left to the model to interpret.

Comments11 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑