arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16806cs.ROcs.AI

当状态成为攻击面:大语言模型驱动的具身智能体中的状态语义注入

Breaking Planner Integrity Boundary: Enviroment State-Text Injection Attack on LLM-Driven Embodied Agents

Jiawei Liu, Jiacheng Guo, Tian Zhang, Yiwei Xu, Juan Wang, Jinlin Fan, Bowen Xiao, Chi Guo, Keyan Guo, Hongxin Hu

首次发表
浏览论文内容

中文总结 AI 辅助

该文聚焦LLM驱动的具身智能体,梳理其技术演进与现有模型,指出其需结合多类信息完成任务落地后执行,为后续状态语义注入攻击研究奠定基础。

中文摘要 AI 辅助

大语言模型(LLMs)已展现出上下文学习、任务分解、逐步推理和代码生成能力,推动其从文本生成模型逐步演变为能够感知环境、调用工具、执行任务的智能体核心。传统LLM智能体通常通过网页、文档、数据库或外部工具获取信息,并根据用户目标生成对应的调用序列;当该技术进一步与机器人系统集成时,大语言模型开始承担任务理解、高层规划和行为决策等功能。SayCan将语言模型的任务推理能力与机器人技能的可供性相结合,Code as Policies和ProgPrompt分别通过策略代码和程序化提示生成机器人任务计划,VoxPoser则利用语言模型和视觉-语言模型构建三维价值图以指导机器人操作[6,7,8,9]。PaLM-E、RT-2和GR00T N1等视觉-语言-动作模型进一步强化了语言、视觉感知与机器人动作之间的联系[10,11,12]。在这类LLM驱动的具身智能体中,模型不仅需要理解用户指令,还需结合场景状态、对象属性、空间关系和执行反馈完成任务 grounding(任务落地),再将生成的动作计划交给技能库、运动规划器或控制器执行。

英文摘要

Large language model (LLM)-driven embodied agents rely on environment states to interpret scenes, generate high-level plans, and drive physical execution, making planner-visible state representations a critical security boundary. Existing attacks primarily manipulate user instructions, prompt contexts, model behavior, or perceptual inputs, while paying limited attention to whether environment-state text itself can serve as deceptive task evidence and propagate beyond planning to affect execution outcomes. Because embodied tasks are constrained by entity grounding, action preconditions, spatial relations, and environmental constraints, planning deviation alone does not guarantee adversarial execution. To address this gap, we investigate environment-state text as an independent attack surface and present the first closed-loop Environment State-Text Injection (ESTI) attack for LLM-driven embodied agents. Without modifying the original user instruction, model parameters, or executor, ESTI reformulates an adversarial objective as false state evidence compatible with the current environment and influences planning and execution through object properties, spatial relations, affordances, task-stage rules, and execution feedback. We further develop ESTI-Bench to evaluate attack propagation across the planning-to-execution closed loop and compare ESTI with Vanilla IPI, EIRAD, and BADROBOT across ProgPrompt/VirtualHome, VoxPoser/RLBench, and AI2-THOR/iTHOR. ESTI consistently outperforms existing baselines, improving planning-level and execution-level attack success rates by up to 89.32\% and 43.69\%, respectively. Further analysis shows that grounding, consistency, and executability jointly determine whether manipulated state evidence can propagate through the embodied closed loop and produce verifiable environmental changes.

发表机构

  • Wuhan University(武汉大学)
  • University at Buffalo(布法罗大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑