arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向世界动作模型的选择性跨视图一致性:无需测试时相机信息的保留视角鲁棒性

Selective Cross-View Consistency for World Action Models: Held-Out Viewpoint Robustness Without Test-Time Camera Information

Bingqi Huang, Bingchuan Wei, Yingkai Cai, Zhaokui Wang

arXiv 2608.21402首次发表:更新:

AI 中文总结

本文针对世界动作模型提出选择性跨视图一致性方法,无需测试时相机信息,在保留视角上提升闭环成功率,同时保持分布内能力,还发现已发表相机鲁棒性数值受腕部相机位姿稳定性混淆。

AI 中文摘要

世界动作模型(World Action Models, WAMs)可联合对未来视频帧与机器人动作进行去噪,其视频先验需具备控制泛化能力,而相机视角变化仍是该模型面临的最具挑战性的扰动轴之一。本文针对此类模型研究了一个特定问题:在使用相同状态的跨视图图像对进行训练时,一致性损失应施加于哪些输出坐标?WAM的去噪目标混合了视图协变坐标(即预测的未来场景)与视图不变坐标(即动作块、未来本体感知及价值)。研究表明,对协变块施加一致性会产生有害影响,会将合法的特定视图内容缩小至其真实值的1/(1+4λ),且在受控实验中验证了该缩小规律。因此,选择性跨视图一致性(Selective Cross-View Consistency, SCVC)仅约束不变块,训练与测试时均无需相机标签、外参、深度或视图合成,且部署接口保持不变。本文在LIBERO-Plus相机轨迹上引入了划分并保留视角的评估协议,该协议将分布匹配的上限与对保留视角的真实插值、外推分离开,采用匹配对训练的控制组将一致性项的影响与对的暴露影响隔离开。在训练范围之外的保留轨道视角上,SCVC相比匹配控制组提升了12.2个点的闭环成功率(95%置信区间[7.4, 17.0];在独立的第二个随机种子下提升15.5个点,置信区间[11.7, 19.4])——该效果在另外两个相机轴上也得到了重复;而在训练范围内的插值中,两个随机种子均未获得增益(分别为-1.2和-4.3个点),且分布内能力得以保留(分别为-0.6和-0.2个点)。本文还报告了跨骨干网络的审计结果,表明已发表的相机鲁棒性数值被腕部相机位姿稳定性所混淆。

英文摘要

World action models (WAMs) jointly denoise future video frames and robot actions, and the video prior is expected to generalize their control. Camera viewpoint change remains one of their hardest perturbation axes. We study a question specific to this model class: when training with same-state cross-view image pairs, on which output coordinates should a consistency loss be imposed? The WAM denoising target mixes view-covariant coordinates, namely the predicted future scene, with view-invariant coordinates, namely the action chunk, future proprioception, and value. We show that consistency applied to the covariant block is provably harmful, shrinking legitimate view-specific content to a fraction $1/(1+4λ)$ of its true value, and we verify this shrinkage law in controlled experiments. Selective cross-view consistency (SCVC) therefore constrains only the invariant block, requires no camera labels, extrinsics, depth, or view synthesis at training or test time, and leaves the deployment interface unchanged. We introduce a carve-and-hold-out evaluation protocol on the LIBERO-Plus camera track that separates a distribution-matched ceiling from genuine interpolation and extrapolation to held-out viewpoints, with a matched pair-trained control isolating the effect of the consistency term from pair exposure. On held-out orbital viewpoints beyond the training envelope, SCVC improves closed-loop success over the matched control by 12.2 points (95% CI [7.4, 17.0]; +15.5, CI [11.7, 19.4], under an independent second seed) -- an effect two further camera axes replicate -- while interpolation within the envelope shows no gain in either seed (-1.2 and -4.3 points) and in-distribution competence is preserved (-0.6, -0.2). We also report a cross-backbone audit showing that published camera-robustness numbers are confounded by wrist-camera pose stability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑