arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TMT:面向未见任务的视觉-语言-动作策略运行时后门检测

TMT: Runtime Backdoor Detection for Vision-Language-Action Policies on Unseen Tasks

Zirun Zhou, Jingfeng Zhang, HaoChuan Xu, Xizhe Zhang, Elliott Wen, Jing Sun, Hong Jia

arXiv 2610.09462首次发表:更新:

发表机构

The University of Auckland; Fudan University; Institute of Science Tokyo(奥克兰大学; 复旦大学; 东京科学大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出TMT,一种基于Token流形和潜在转移建模的运行时后门检测器,用于在未见任务上检测视觉-语言-动作策略中的后门激活,并通过自蒸馏实现策略净化,达到最先进的检测性能。

AI 中文摘要

被植入后门的视觉-语言-动作(VLA)策略在触发器出现时,既能保持良性任务性能,又能产生恶意动作。检测此类激活十分困难,因为恶意行为可能由各自看似合理的动作组成,而未见任务会引入观测和行为上的合理变化。我们提出TMT,一种基于Token流形和潜在转移建模的运行时后门检测器。该方法在良性轨迹上训练,其两个分支分别评估输入令牌结构和相邻层潜在动态中的预测误差。由令牌流形分支识别出的可疑轨迹,经潜在偏差确认后,将指导后续监控的转移选择。我们进一步探索通过自蒸馏进行策略净化:后门策略的冻结副本提供良性输入动作,以监督学生模型在配对的良性和触发观测上的学习,无需单独的干净参考策略。为进行评估,我们改编了传统后门检测器,并将异常和故障检测方法重新用作VLA后门检测器。在与十个基线的后验比较中,TMT在三种VLA后门攻击下的未见任务上达到了最先进的后门检测性能。我们的项目页面可在该https URL获取。

英文摘要

Backdoored vision-language-action (VLA) policies can preserve benign task performance while producing malicious actions when a trigger appears. Detecting such activation is difficult because malicious behavior can comprise individually plausible actions, while unfamiliar tasks introduce legitimate changes in observations and behavior. We introduce TMT, a runtime backdoor detector based on Token Manifold and latent Transition modeling. Trained on benign rollouts, its two branches assess input-token structure and prediction errors in adjacent-layer latent dynamics. A suspicious rollout identified by the token manifold branch, once confirmed through latent deviations, guides transition selection for subsequent monitoring. We further explore policy purification through self-distillation: a frozen copy of the backdoored policy provides benign-input actions to supervise a student on paired benign and triggered observations, without requiring a separate clean reference policy. For evaluation, we adapt traditional backdoor detectors and repurpose anomaly and failure detection methods as VLA backdoor detectors. In a post-hoc comparison with ten baselines, TMT achieves state-of-the-art backdoor detection performance on unseen tasks across three VLA backdoor attacks. Our project page is available at https://zzr42.github.io/tmt/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑