arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在变化物体布局中的终身小物体导航:一个基准与方法

Lifelong small-object navigation in changing object layouts: a benchmark and method

Jiagan Huang, Zikun Zhou, Zijian Ni, Hongpeng Wang, Guangming Lu, Jun Yu, Wenjie Pei

arXiv 2610.10125首次发表:更新:

发表机构

Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对家用机器人在变化布局中持续导航至小物体的挑战,本文提出LiSoNav-COL任务、LiSoNav-Eval基准及IVAM-Nav方法,通过多视角观察与视角锚定记忆实现高效定位与适应重定位。

AI 中文摘要

家用机器人需要持续在同一环境中导航到不同的物体,其中许多物体是小型且便携的,例如工具和玩具。它们较小的视觉足迹和频繁的遮挡使得可靠观察变得困难,并且它们可能被人在机器人未观察到变化的情况下移动。我们将这一具有挑战性的任务定义为在变化物体布局中的终身小物体导航(LiSoNav-COL)。智能体必须寻找合适的视角进行可靠观察,积累并重用场景知识以高效定位后续目标,并在物体重新定位后更新过时的记忆。为了消除对先前场景扫描的需求,我们还要求智能体以空场景记忆开始导航。尽管实用,但该任务仍缺乏围绕其定义假设设计的基准。为填补这一空白,我们引入了LiSoNav-Eval,一个专门的基准,涵盖28个室内场景和45个小物体类别。其终身导航序列包括未改变和重新定位的目标,以评估记忆重用和对物体重新定位的适应性。为解决这一具有挑战性的任务,我们提出了一种基于多视角检查与视角锚定记忆的导航方法,称为IVAM-Nav。IVAM-Nav主动从互补视角观察支撑表面,以实现可靠的小物体感知,并将其产生的记忆锚定到观察视角,支持在相似观察条件下的关系记忆重用和重新验证。在LiSoNav-Eval上的大量实验表明,IVAM-Nav相对于代表性方法具有优越性能。基准分析还表明,较小的物体、较大的环境和较长的重新定位距离构成更大的挑战。数据集和代码可在此处获取。

英文摘要

Household robots need to continually navigate to different objects in the same environment, many of which are small and portable, such as tools and toys. Their small visual footprint and frequent occlusion make reliable observation difficult, and they may be moved by people without the robot observing the changes. We formulate this challenging task as Lifelong Small-object Navigation in Changing Object Layouts (LiSoNav-COL). Agents must seek suitable viewpoints for reliable observation, accumulate and reuse scene knowledge to efficiently locate subsequent targets, and update outdated memory after object relocation. To eliminate the need for prior scene scanning, we also require agents to start navigation with empty scene memory. Although practical, this task still lacks benchmarks designed around its defining assumptions. To bridge this gap, we introduce LiSoNav-Eval, a dedicated benchmark spanning 28 indoor scenes with 45 small-object categories. Its lifelong navigation sequences include both unchanged and relocated targets to evaluate memory reuse and adaptation to object relocation. To address this challenging task, we propose a navigation method based on multi-view Inspection with Viewpoint-Anchored Memory, dubbed IVAM-Nav. IVAM-Nav actively observes supporting surfaces from complementary viewpoints for reliable small-object perception and anchors the resulting memory to their observation viewpoints, supporting relational memory reuse and revalidation under similar viewing conditions. Extensive experiments on LiSoNav-Eval demonstrate favorable performance of IVAM-Nav against representative methods. Benchmark analyses also show that smaller objects, larger environments, and longer relocation distances pose greater challenges. The dataset and code are available here.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑