arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26184cs.ROcs.AIcs.CR

无声破坏:基于内部状态触发的LLM驱动机器人系统后门攻击

Silent Sabotage: Internal State Triggered Backdoor Attacks on LLM-Powered Robotic Systems

Doniyorkhon Obidov, Shivayogi Akki, Tan Chen, Kaichen Yang

首次发表
浏览论文内容

中文总结 AI 辅助

本文首次系统研究LLM驱动机器人系统中基于历史记录的后门攻击,通过操纵指令植入由机器人自身动作序列触发的隐蔽后门,实验证明攻击成功率近乎完美且极难检测,揭示了内部状态安全漏洞的紧迫性。

中文摘要 AI 辅助

大型语言模型(LLMs)与机器人控制系统的集成正在催生新一代具备复杂推理与规划能力的自主智能体。尽管这一范式转变加速了技术进步,但也引入了尚未被充分探索的新型安全风险。当前针对LLM后门的研究主要聚焦于由外部刺激(如特定词语、视觉对象或环境状态)触发的攻击。这些攻击虽然威力强大,却忽视了一类更为隐蔽的漏洞,其触发条件内嵌于智能体自身的运行逻辑之中。本文首次对基于历史记录的后门攻击在LLM驱动的机器人系统中进行了全面研究。我们证明,攻击者可通过操纵指令,将隐蔽后门植入基于LLM的机器人控制器中。该后门并非由外部线索触发,而是由机器人自身过去动作的特定罕见序列激活。它在正常运行期间保持休眠状态,不影响机器人的正常功能,但可被激活以诱导恶意行为,如完全停止或碰撞。我们在包含多种机器人和LLM的模拟环境中进行的实验表明,这种基于历史记录的攻击具有极高的有效性,实现了近乎完美的攻击成功率,同时极难被检测。这些发现揭示了自主系统中一个关键且此前未被解决的漏洞,并强调了针对智能体内部状态采取安全措施的紧迫性。

英文摘要

The integration of Large Language Models (LLMs) into robotic control systems is enabling a new generation of autonomous agents capable of complex reasoning and planning. While this paradigm shift accelerates progress, it also introduces novel security risks that remain largely unexplored. Current research into LLM backdoors has focused on attacks triggered by external stimuli, such as specific words, visual objects, or environmental states. These attacks, while potent, overlook a more insidious class of vulnerability where the trigger is internal to the agent's own operational logic. This paper presents the first comprehensive study of history-based backdoor attacks on LLM-powered robotic systems. We demonstrate that an attacker can embed a stealthy backdoor into an LLM-based robot controller by manipulating its instructions. This backdoor is triggered not by an external cue, but by a specific, rare sequence of the robot's own past actions. It remains dormant during normal operation, preserving the robot's utility, but can be activated to induce a malicious behavior, such as a complete stop or a collision. Our experiments, conducted in a simulated environment with a variety of robots and LLMs, show that this history-based attack is highly effective, achieving a near-perfect attack success rate while remaining exceptionally difficult to detect. These findings reveal a critical and previously unaddressed vulnerability in autonomous systems and underscore the urgent need for security measures that account for an agent's internal state.

发表机构

  • Michigan Technological University(密歇根理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑