arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.29169cs.ROcs.AI

ActFovea:基于时空视觉-动作一致性的VLA策略运行时安全保障

ActFovea: Runtime Safeguarding for VLA Policies via Spatiotemporal Visual-Action Consistency

Wenda Yu, Tianshi Wang, Fengling Li, Xin Li, Jingjing Li, Lei Zhu

AI总结:

ActFovea是一种即插即用的VLA策略运行时安全保障框架,通过构建动作条件化凹形区域、检测时空一致性及触发安全故障,提升了机器人操纵在各类干扰下的成功率与安全性。

AI中文摘要:

视觉-动作(VLA)策略在机器人操纵中表现出色,但仍易受运行时干扰影响,破坏视觉观测、机器人状态与执行动作间的时间对齐。本文提出ActFovea,一种即插即用的安全保障框架,无需重新训练或修改底层VLA策略即可检测并缓解此类故障。ActFovea利用机器人运动学、本体感受状态及近期动作构建动作条件化的凹形区域,保留与接触相关的区域和预测运动走廊,同时抑制与任务无关的视觉内容。它通过评估视觉运动和观测新鲜度是否与几何、本体感受及动作转换保持一致来检测运行时风险。对于可恢复干扰,ActFovea构建特定于干扰的候选观测,并在验证生成的动作块后才接受恢复;当陈旧或重放的观测使可靠恢复不可能时,它会调用有界安全故障程序。在多个LIBERO套件上对π₀的闭环评估中,ActFovea在局部视觉覆盖下的成功率从49.3%提升至90.3%,缩小了与干净性能差距的93.7%;还分别将动作漂移和视觉延迟下的成功率提高了7.0和9.8个百分点,同时保留了干净任务性能;在冻结观测重放场景中,ActFovea在所有试验中都及时触发安全故障,无未受保护的故障。这些结果表明,时空视觉-动作一致性为VLA策略的运行时安全保障提供了有效基础。

英文摘要:

Vision-language-action (VLA) policies achieve strong performance in robotic manipulation but remain vulnerable to runtime disturbances that break the temporal alignment among visual observations, robot states, and executed actions. We introduce ActFovea, a plug-and-play safeguarding framework that detects and mitigates such failures without retraining or modifying the underlying VLA policy. ActFovea uses robot kinematics, proprioceptive states, and recent actions to construct action-conditioned foveated regions that retain contact-relevant areas and predicted motion corridors while suppressing task-irrelevant visual content. It detects runtime risks by evaluating whether visual motion and observation freshness remain consistent with geometric, proprioceptive, and action transitions. For recoverable disturbances, ActFovea constructs disturbance-specific candidate observations and accepts a recovery only after verifying the resulting action chunk. When stale or replayed observations make reliable recovery impossible, it invokes a bounded safe-failure procedure. In closed-loop evaluations of $π_0$ across multiple LIBERO suites, ActFovea increases success under localized visual overlays from 49.3\% to 90.3\%, closing 93.7\% of the gap to clean performance. It further improves success under action drift and visual delay by 7.0 and 9.8 percentage points, respectively, while preserving clean-task performance. Under frozen-observation replay, ActFovea triggers timely safe failure in all trials, with no unprotected failures. These results demonstrate that spatiotemporal visual-action consistency provides an effective basis for runtime safeguarding of VLA policies.

补充信息

↑