发表机构
Cornell University(康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究利用机制可解释性探究VLA的π₀.₅残差流,发现任务进度可从激活值线性读取,基于此构建的探测器可作为无标签OOD检测器,性能接近最先进方法,为监控部署的视觉运动策略提供轻量可解释路径。
AI 中文摘要
视觉-语言-动作模型(VLAs)正快速朝着作为通用操作策略部署的方向发展,但目前我们缺乏理解这些模型内部表征内容或在运行时对其进行监控的基础工具。利用机制可解释性的思路,我们探究了π₀.₅的残差流,发现任务进度(即轨迹中剩余的归一化时间)可从激活值中线性读取。我们发现该信号存在于预训练的PaliGemma主干网络中,且在针对任何机器人特定数据进行训练前就已存在。在使用多提示数据进行训练时,单个线性探测器可泛化到未见过的任务,并在语言反事实条件下发生变化,但无法实现对策略的有效控制。这些特性使该信号可直接用于为已部署的VLAs添加监控工具。我们将该探测器用作简单的无标签分布外(OOD)检测器,以检测停滞的任务进度,发现其性能与最先进方法相当。我们的结果表明,VLAs具有丰富的、可线性读取的语义量内部表征,例如任务进度,且学习读取这些信号为监控已部署的视觉运动策略提供了一条轻量、可解释的路径。
英文摘要
Vision-language-action models (VLAs) are moving rapidly towards deployment as general-purpose manipulation policies, but we currently lack basic tools for understanding what these models represent internally or for monitoring them at runtime. Leveraging ideas from mechanistic interpretability, we probe the residual stream of $π_{0.5}$ and find that task progress, the normalized time remaining in a trajectory, is linearly readable from the activations. We find that this signal is present in the pretrained PaliGemma backbone prior to training on any robot-specific data. A single linear probe generalizes to unseen tasks and varies under language counterfactuals when trained on multi-prompt data, but does not enable meaningful steering of the policy. These properties make the signal directly useful for instrumenting deployed VLAs. We use the probe as a simple label-free OOD detector, which detects stalled task progress, and find it competitive with state-of-the-art methods. Our results suggest that VLAs have rich, linearly readable internal representations of semantic quantities like task progress, and that learning to read these signals offers a lightweight, interpretable path toward monitoring deployed visuomotor policies.