TrapVLA:将视觉-语言-动作模型困在配置好的故障模式中
TrapVLA: Trapping Vision-Language-Action Models in Configured Failure Modes
浏览论文内容
中文总结 AI 辅助
本研究针对VLA模型提出名为TrapVLA的后门攻击方法,通过配置故障诱捕任务诱导模型进入指定故障模式,构建了相关基准并验证了攻击效果。
中文摘要 AI 辅助
本研究提出了一种针对视觉-语言-动作(Vision-Language-Action, VLA)模型的新型后门攻击任务——配置故障诱捕(Configured Failure Trapping),该任务旨在通过隐蔽的文本触发器激活攻击并诱导模型进入配置好的故障模式。与以往将任何任务失败都视为攻击成功的后门攻击不同,配置故障诱捕要求攻击者控制机器人的失败方式(例如,使机器人以指定的位置偏移进行抓取),这使得该任务更具挑战性且更难被检测。为支持这一新任务,我们设计了一个用于合成高质量目标轨迹的有效数据引擎,以及一套用于测量配置故障保真度的自动化工具。在此基础上,我们构建了两个新基准,即Trap-LIBERO和Trap-RoboTwin,它们在四种代表性故障模式下实例化了配置故障诱捕任务。为解决该任务,我们发现稀疏动作偏差是一个关键挑战,因此提出了一种名为TrapVLA的新方法,该方法明确学习触发器诱导的动作残差,以引导策略走向配置好的故障行为。在模拟基准和真实机器人环境中进行的大量实验表明,TrapVLA能有效将配置好的故障模式注入VLA模型,同时在干净数据上基本保持原有性能。项目页面:this https URL
英文摘要
This work introduces Configured Failure Trapping, a novel backdoor attack task against Vision-Language-Action (VLA) models, which aims to activate attacks through stealthy textual triggers and induce configured failure modes. Unlike prior backdoor attacks that treat any task failure as a successful attack, Configured Failure Trapping requires the attacker to control how the robot fails (e.g., causing the robot to grasp with a specified positional offset), making it substantially more challenging and hard to detect. To support the new task, we propose an effective data engine for synthesizing high-quality target trajectories and an automated suite for measuring configured-failure fidelity. Then, based on this foundation, we construct two new benchmarks, namely Trap-LIBERO and Trap-RoboTwin, that instantiate Configured Failure Trapping across four representative failure modes. To address this task, we identify sparse action deviation as a critical challenge and accordingly propose a novel method named TrapVLA, which explicitly learns trigger-induced action residuals to steer the policy toward the configured failure behavior. Extensive experiments across simulation benchmarks and real-world robotic settings show that TrapVLA effectively injects configured failure modes into VLA models while largely preserving performance on clean data. Project page: https://john-liua.github.io/TrapVLA/
发表机构
- School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)
- Pengcheng Laboratory(鹏城实验室)
- The University of Hong Kong(香港大学)
- Jiangxing Intelligence (Guizhou) Technology Inc.(匠星智能(贵州)科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。