arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

更好的想象式回滚是否意味着更好的机器人控制?反馈下世界模型评估的对照研究

Do Better Imagined Rollouts Mean Better Robot Control? A Controlled Study of World-Model Evaluation Under Feedback

Dharini Raghavan, Amritpal Singh

arXiv 2609.02811首次发表:更新:

发表机构

Georgia Institute of Technology; Emory University(佐治亚理工学院; 埃默里大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究通过差速驱动路径跟踪任务对照实验发现,离线回滚的感知与修正模式需匹配闭环操作,轨迹重放比回滚误差更能反映机器人控制性能,评估预测模型需明确预测时间范围和测量更新计划

AI 中文摘要

预测模型在机器人技术中越来越多地用于状态估计、规划、控制和策略评估,但它们通常是根据固定时间范围内的开环预测精度来评判的。在闭环操作中,机器人会反复执行动作、接收新的测量值、更新状态估计并重新计算控制。我们在带有偏置里程计和间歇性地标感知的差速驱动路径跟踪任务中研究了这种差异。我们使用轨迹重放、20步无测量回滚和闭环跟踪,在24种感知条件下评估了6种状态估计器。与回滚误差相比,重放位置RMSE与闭环横向RMSE的相关性更强(Spearman秩相关系数分别为0.923和0.774);在5/24的条件下,重放选择的估计器不同于闭环最优估计器,而回滚指标的这一比例为18/24。随后,我们改变回滚时间范围和测量更新间隔。当H=20时,秩一致性从每次步骤都有测量时的rho=0.916降至无测量时的rho=0.774。时间范围-更新网格显示,在保留定期修正的情况下,长预测时间范围仍具有信息性,而没有修正的长回滚可能会产生与闭环行为显著不同的排名。我们还测试了针对更长感知中断训练的循环估计器,这在感知退化组合情况下改善了EKF锚定模型,将GRU-EKF的横向RMSE从1.72米降至1.06米,但该增益在单独中断或估计器架构之间并不一致。这些结果表明,机器人技术中的预测模型评估应同时指定预测时间范围和测量更新计划;对于用于反馈的模型,离线回滚在其感知和修正模式反映闭环操作时最具信息性。代码可在该https URL获取

英文摘要

Predictive models are increasingly used in robotics for state estimation, planning, control, and policy evaluation, yet they are often judged by open-loop prediction accuracy over a fixed horizon. In closed-loop operation, a robot repeatedly acts, receives new measurements, updates its state estimate, and recomputes control. We study this difference in a differential-drive path-tracking task with biased odometry and intermittent landmark sensing. Six state estimators are evaluated across 24 sensing conditions using trajectory replay, a 20-step measurement-free rollout, and closed-loop tracking. Replay position RMSE correlates more strongly with closed-loop cross-track RMSE than rollout error (Spearman rho = 0.923 vs. 0.774) and selects a different estimator from the closed-loop optimum in 5/24 conditions, compared with 18/24 for the rollout metric. We then vary rollout horizon and measurement-update interval. With H=20, rank agreement decreases from rho = 0.916 with measurements at every step to rho = 0.774 with no measurements. A horizon-update grid shows that long prediction horizons remain informative when regular corrections are retained, whereas long rollouts without correction can produce rankings that differ substantially from closed-loop behavior. We also test recurrent estimators trained on longer sensing outages. This improves the EKF-anchored models under combined sensing degradation, reducing GRU-EKF cross-track RMSE from 1.72 m to 1.06 m, but the gain is not consistent across isolated outages or estimator architectures. These results show that predictive-model evaluation in robotics should specify both prediction horizon and measurement-update schedule. For models used in feedback, offline rollouts are most informative when their sensing and correction pattern reflects closed-loop operation. Code is available at https://github.com/rdharini2001/Robot_World_Model

Comments20 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑