基于集合训练的视觉运动策略的概率可达动作验证
Probabilistic Reachable-Action Verification of Visuomotor Policies via Set-Based Training
AI总结:
本文针对视觉运动策略可达性分析的高成本与保守性问题,提出冻结视觉编码器、在低维接口做集合传播的集合训练方法,经实验验证该方法可减小概率可达动作半径并保留任务能力。
AI中文摘要:
视觉运动策略的可达性分析难度在于,大型视觉编码器使端到端集合传播的计算成本过高且过于保守。因此,本文冻结视觉编码器,将集合传播限制在其与下游策略之间的低维接口,该接口集合由预留的相机位姿扰动校准。通过 zonotope 将此集合经策略传播,得到终端输出包围盒宽度,集合训练直接对其优化。评估阶段,从规定分布采样相机位姿扰动,采用 rollout 级拆分共形校准将所得动作偏差分数转换为具有有限样本覆盖率的概率可达动作半径。在受控操纵实验中,集合训练减小了该半径,同时保留闭环任务能力,而匹配的仅行为、观测一致性及逐点对抗对照组均产生更大半径。
英文摘要:
Reachability analysis for visuomotor policies is difficult because large visual encoders make end-to-end set propagation computationally expensive and excessively conservative. We therefore freeze the visual encoder and confine set propagation to a low-dimensional interface between it and the downstream policy, with the interface set calibrated from held-out camera-pose perturbations. Propagating this set through the policy with zonotopes yields a terminal output-enclosure width that set-based training optimizes directly. During evaluation, camera-pose perturbations are sampled from the prescribed distribution, and rollout-level split conformal calibration converts the resulting action-deviation scores into a probabilistic reachable-action radius with finite-sample coverage. In controlled manipulation experiments, set-based training reduces this radius while preserving closed-loop task capability, and matched behavior-only, observational-consistency, and pointwise-adversarial controls all leave a larger radius.