AuthGuard-R:面向LLM控制机器人的安全合规任务劫持与双门防御
AuthGuard-R: Safety-Compliant Mission Hijacking and Dual-Gate Defense for LLM-Controlled Robots
浏览论文内容
中文总结 AI 辅助
针对LLM控制机器人可能被物理安全但违反用户授权任务劫持的问题,提出MissionPAIR攻击框架和AuthGuard-R双门防御授权层,通过形式化验证和实验证明其能有效阻止未授权动作。
中文摘要 AI 辅助
大型语言模型越来越多地被用作移动机器人、机器人操作臂和自动驾驶汽车的高层规划器。近期研究表明,这些系统可能受到恶意文本、语音、视觉指令、检索文档和中毒感知上下文的影响。大多数防御机制只询问所提议的动作在物理上是否安全。本文研究一个不同的问题:一个动作可能在物理上安全,但仍然违反用户授权的任务。攻击者可能使配送机器人偏离路线、替换已批准的对象、扩展机器人的操作区域、激活不必要的传感器或延迟任务,而不会造成即时的物理危险。我们将这种攻击称为“安全合规任务劫持”。我们提出MissionPAIR,一种自适应攻击框架,用于搜索能够通过安全门但违反已认证任务的可执行计划。我们还提出AuthGuard-R,一个确定性的授权层,将每个可执行动作绑定到已签名的任务、机器人身份、对象和区域范围、当前状态、时间和输入来源。AuthGuard-R与独立的安全门协同工作,形成双门架构。我们形式化定义了基于机器人轨迹的任务策略,定义了安全游戏,并证明了授权健全性、任务非升级性、重放抵抗性、机器人绑定、来源分离、阈值批准安全性、审计日志防篡改性和轨迹级组合性。我们报告了与Claude Haiku 4.5和开源Qwen2.5 7B规划器的初步跨模型评估。在240次实时攻击试验中,规划器在109次试验中遵循了注入的任务偏差;AuthGuard-R拒绝了所有109次由此产生的未授权动作。另外手工构建的包含十一个协议级和策略级攻击的测试套件也被完全阻止。
英文摘要
Large language models are increasingly used as high-level planners for mobile robots, robot manipulators, and autonomous vehicles. Recent studies show that these systems can be influenced through malicious text, speech, visual instructions, retrieved documents, and poisoned sensory context. Most defenses ask whether a proposed action is physically safe. This paper studies a different problem: an action may be physically safe and still violate the mission authorized by the user. An attacker may redirect a delivery robot, replace an approved object, extend a robot's operating region, activate an unnecessary sensor, or delay a mission without creating an immediate physical hazard. We call this attack \emph{safety-compliant mission hijacking}. We propose MissionPAIR, an adaptive attack framework that searches for executable plans that pass a safety gate while violating an authenticated mission. We also propose AuthGuard-R, a deterministic authorization layer that binds every executable action to a signed mission, robot identity, object and region scope, current state, time, and input provenance. AuthGuard-R operates with an independent safety gate, giving a dual-gate architecture. We formalize mission policies over robot traces, define security games, and prove authorization soundness, mission non-escalation, replay resistance, robot binding, provenance separation, threshold-approval security, audit-log tamper evidence, and trace-level composition. We report a preliminary cross-model evaluation with Claude Haiku~4.5 and the open-source Qwen2.5~7B planner. Across 240 live attack trials, the planners followed an injected mission deviation in 109 trials; AuthGuard-R rejected all 109 resulting unauthorized actions. A separate hand-constructed suite of eleven protocol- and policy-level attacks was also blocked completely.
发表机构
- National Institute of Technology Warangal(瓦朗加尔国立理工学院)
机构由 AI 辅助整理,请以论文原文为准。