发表机构
Seoul National University(首尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究分析VLA模型在相机黑屏和冻结故障下的物理失效模式,发现本体感觉可部分补偿,并评估两种缓解方法,旨在利用可用信息限制危险运动。
AI 中文摘要
不可靠的视觉输入会损害视觉-语言-动作(VLA)模型的任务性能,并可能带来潜在的身体安全风险。我们分析了π0.5和GR00T模型在图像黑屏和冻结等输入故障下的行为。我们发现,即使任务成功率同样较低,黑屏和冻结也会产生不同的物理失效模式:冻结会导致更极端的关节行为,而夹爪闭合后的黑屏可能导致更多物体掉落,尤其是在没有本体感觉的情况下。选择性干预研究表明,本体感觉(当前机器人状态)部分补偿了被移除的机器人图像,并减少了非目标接触。然而,当腕部视角的物体信息被移除时,即使有剩余场景视图的辅助,本体感觉也无法充分恢复任务成功率。随后,我们评估了两种缓解方法:相机黑屏训练和无训练替换故障视觉嵌入。两者在特定条件下均能提高任务成功率,但可能增加意外接触或对周围物体的干扰。真实机器人试验进一步表明,在相机故障下成功执行仍可能涉及意外的物理交互。这些发现促使设计VLA策略,利用相机故障下仍可用的机器人和物体信息来限制危险运动。
英文摘要
Unreliable visual inputs can harm task performance and cause potential physical safety risks for vision-language-action (VLA) models. We analyze how $π0.5$ and GR00T models act under input faults such as image blackouts and freezing. We find that blackout and freezing produce distinct physical failure modes even when task-success rates are similarly low: freezing causes more extreme joint behavior, whereas blackout after gripper closure can cause more object drops, most markedly without proprioception. Selective intervention studies reveal that proprioception (current robot state) partly compensates for the removed robot depictions and reduces non-target contact. However, it cannot sufficiently restore task success when wrist-view object information is removed, even when aided by the remaining scene view. We then evaluate two mitigation approaches: camera-blackout training and training-free replacement of faulty visual embeddings. Both improve task success in selected conditions, but can increase unintended contact or disturbance to surrounding objects. Real-robot trials further show that successful execution under camera faults can still involve unintended physical interactions. These findings motivate designing VLA policies that use the robot and object information still available under camera faults to limit hazardous motion.