arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29942cs.CRcs.AI

影响并非权威:在使用工具的大语言模型智能体中,因果护栏信号为何会让合法工具使用看起来像攻击

Influence Is Not Authority: When Causal Guardrail Signals Make Legitimate Tool Use Look Like an Attack in Tool-Using LLM Agents

Tanzim Ahad, Ismail Hossain, Md Jahangir Alam, Sai Puppala, Syed Bahauddin Alam, Sajedul Talukder

首次发表
浏览论文内容

中文总结 AI 辅助

该研究指出基于影响的大语言模型智能体护栏无法可靠区分合法与恶意工具操作,经多组实验验证了该局限,并测试了语义监控器、影子护栏等方案的表现,为安全决策提供了参考。

中文摘要 AI 辅助

当前最先进的基于影响的护栏存在一个关键局限:当合法的、用户授权的操作与恶意的、未授权的操作都依赖外部工具信息时,它们无法可靠地区分这两种操作。这种模糊性会导致良性操作触发不必要的验证和干预,降低效用并增加延迟。我们通过对24个基础案例衍生出的96种条件进行授权等价审计,揭示了这一局限。在匹配的源比较中,我们保持授权、确切执行的操作及其预期效果固定,仅改变所需值的来源——是来自用户还是合法工具的结果。尽管操作本身未变,但这种无害的位置变更在Llama和Gemma评分器下的全部24个案例中,都将因果信号转向了攻击区域。匹配的未授权对照实验表明,信号仍对攻击敏感,但这种良性位置变更产生的平均分数偏移,比授权的实际变化更大。架构级评估显示,这种不匹配如何在护栏设计中传播。使用语义监控器时,攻击成功率为0%,效用为28%;而不使用语义监控器时,攻击成功率为16%,效用为60%。基于影子的护栏允许所有测试的良性运行,但并未更频繁地拒绝匹配的未授权操作:57.5%的未授权运行在到达后续安全检查前自动通过,而授权运行的这一比例为29.2%。这些结果表明,所研究的因果信号揭示了塑造操作的因素,但无法可靠地编码该操作是否被授权;参考构建与路由是有效安全决策的组成部分。

英文摘要

The key limitation of current state-of-the-art influence-based guardrails is that they do not reliably distinguish a legitimate, user-authorized action from a malicious, unauthorized action when both rely on external tool information. This ambiguity can cause benign actions to trigger unnecessary verification and intervention, reducing utility and adding latency. We expose this limitation through an authorization-equivalence audit of 96 conditions derived from 24 base cases. Within matched source comparisons, we hold authorization, the exact committed action, and its intended effect fixed, changing only whether a required value comes from the user or a legitimate tool result. Although the action remains unchanged, this harmless relocation shifts the causal signal toward the attack region in all 24 cases under both Llama and Gemma scorers. Matched unauthorized controls show that the signal remains attack-sensitive, yet the benign relocation produces a larger average score shift than the actual change in authorization. Architecture-level evaluation shows how this mismatch propagates through guardrail designs. With a semantic monitor, attack success is 0% and utility is 28%, compared with 16% and 60% without it. A shadow-based guardrail allows every tested harmless run, yet does not reject matched unauthorized actions more often overall: 57.5% of unauthorized runs pass automatically before reaching the later security check, compared with 29.2% of authorized runs. These results show that the studied causal signal reveals what shaped an action without reliably encoding whether the action was authorized, and that reference construction and routing are integral to the effective security decision.

发表机构

  • University of Texas at El Paso(德克萨斯大学埃尔帕索分校)
  • New Mexico Tech(新墨西哥理工学院)
  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

↑