arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于强化学习的含噪定位下的宏动作拓扑导航

Macro-Action Topological Navigation under Noisy Localization using Reinforcement Learning

Simon Hakenes, Tobias Glasmachers

arXiv 2608.23055首次发表:更新:

发表机构

Institute of Neural Computation, Ruhr University Bochum(鲁尔大学波鸿分校神经计算研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出一种基于强化学习的含噪定位下的宏动作拓扑导航方法,通过以物体为中心的位姿估计替代真实位姿,在Habitat模拟器中实现智能体仅用视觉完成3D公寓内的导航任务。

AI 中文摘要

从原始像素在大型、照片级真实感的3D公寓中导航,被广泛认为是普通强化学习无法实现的。我们构建了一个智能体,它仅通过相机就能完成该任务,同时估计自身的位姿。该智能体需要按顺序到达多个目标物体,且目标物体的位置在不同回合间会发生变化,因此它必须进行探索以找到这些物体。该智能体基于我们早期以物体为中心的拓扑控制器构建,早期控制器仍从模拟器中读取智能体的真实位姿及其物体检测结果。在此,我们将该真实位姿替换为机载的、以物体为中心的估计值。对于每个物体,我们维护一组ORB特征,当再次看到该物体时,这些特征会生成粗略的位姿测量值,一个极简扩展卡尔曼滤波器(EKF)会将该测量值与运动模型进行融合。与真实机器人的情况类似,执行的运动存在噪声,位姿估计会发生漂移,但智能体与附近物体会一同漂移,因此局部一致的位姿足以让智能体跟随每条短边,随后在视觉上对准目标,这使我们能用一个小得多的模型替代完整的SLAM,更接近生物导航的运作方式。在照片级真实感的Habitat模拟器中,该智能体仅通过视觉就能到达目标物体,且其位姿仅需满足局部一致性即可。

英文摘要

Navigating large, photorealistic 3D apartments from raw pixels is widely considered infeasible for plain reinforcement learning. We build an agent that does it anyway, estimating its own pose from the camera alone. The agent has to reach several target objects in sequence, and their positions change between episodes, so it must explore to find them. It builds on our earlier object-centric topological controller, which still read the agent's true pose and its object detections from the simulator. Here we replace that true pose with an onboard, object-centric estimate. For each object we keep a bank of ORB features that, when the object is seen again, yield a rough pose measurement, which a minimal Extended Kalman Filter (EKF) fuses with a motion model. As on a real robot, the executed motions are noisy. The estimate drifts, but the agent and the nearby objects drift together, so a locally consistent pose is enough to follow each short edge and then home in visually on the target, which lets us replace full SLAM with a much smaller model, closer to how biological navigation appears to work. In the photorealistic Habitat simulator, the agent reaches its target objects from vision alone, with a pose that only needs to be locally consistent.

Comments15 pages, Accepted at the Artificial Intelligence Symposium (AIS) 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑