发表机构
China University of Mining and Technology; Shanghai Jiao Tong University; Shanghai Innovation Institute(中国矿业大学; 上海交通大学; 上海创新研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出WM-VS框架,通过进度对齐世界模型解决闭环视觉伺服中预测-控制不匹配问题,在真实7自由度系统上实现高成功率误差收缩,并展现良好泛化能力。
AI 中文摘要
闭环视觉伺服需要能够指示动作是否减少任务误差的预测,而不仅仅是动作是否合理。我们将这一差距称为预测-控制不匹配,并引入WM-VS,一种面向闭环视觉伺服的目标中心、进度对齐的世界模型框架。离线目标区域DINOv2对应关系定义了用于平移、缩放和面内旋转的带符号四维伺服坐标。第一阶段将动作条件潜在状态转移与该坐标对齐;第二阶段冻结世界模型,并通过动作模仿、后果监督以及有利于误差收缩的短想象展开来训练反应式关节速度策略。部署仅依赖RGB且具有反应性,无需在线轨迹优化。在真实的7自由度眼对手系统上,WM-VS在30/30次试验中达到转角均方根误差不超过其初始值10%的水平,并在25/30次试验(83.33%)中于最终有效帧保持该标准。移除未来误差对齐会使保持率降至26.67%。学习到的进度信号与未用于训练或控制的外部AprilTag转角误差一致(平均Spearman rho = 0.8778)。无需重新训练,两个未见过的3D目标实现了86.48%和90.27%的平移误差减少,以及70.01%和65.70%的旋转误差减少。这些结果将进度对齐的动作后果与重复闭环校正和迁移联系起来。代码和数据将作为开源发布。
英文摘要
Closed-loop visual servoing requires predictions that indicate whether an action reduces task error, not only whether the action is plausible. We call this gap the prediction-control mismatch and introduce WM-VS, a target-centric progress-aligned world-model framework for closed-loop visual servoing. Offline target-region DINOv2 correspondences define a signed four-dimensional servo coordinate for translation, scale, and in-plane rotation. Stage 1 aligns action-conditioned latent transitions with this coordinate; Stage 2 freezes the world model and trains a reactive joint-velocity policy with action imitation, consequence supervision, and short imagined rollouts that favor error contraction. Deployment is RGB-only and reactive, without online trajectory optimization. On a real 7-DoF eye-to-hand system, WM-VS reaches a corner RMSE no larger than 10 percent of its initial value in 30/30 trials and retains this criterion at the final valid frame in 25/30 (83.33 percent). Removing future-error alignment reduces retention to 26.67 percent. The learned progress signal agrees with an external AprilTag corner error not used for training or control (mean Spearman rho = 0.8778). Without retraining, two unseen 3D targets achieve translation-error reductions of 86.48 percent and 90.27 percent and rotation-error reductions of 70.01 percent and 65.70 percent. These results link progress-aligned action consequences to repeated closed-loop correction and transfer. Code and data will be released as open source.