arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.09267cs.CV

REMIND:用于室内导航的带记忆重新识别

REMIND: RE-Identification with Memory for INDoor Navigation

Pablo Diaz-Pereda, Alejandro Rodriguez-Ramos, David Perez-Saura, Pascual Campoy

首次发表
浏览论文内容

中文总结 AI 辅助

研究室内移动机器人物体重识别难题,提出REMIND在线跟踪器,结合多种技术,在专用室内数据集和ScanNet++上表现出色,IDF1高,完成场景能力强,且系统、框架和数据集已公开。

中文摘要 AI 辅助

室内运行的移动机器人在长时间间隔、显著视角变化和严重光照变化后,必须重新识别先前观察到的物体。这仍然是一个具有挑战性的问题:多目标跟踪方法针对行人与车辆的短期关联进行了优化,行人与车辆重新识别方法缺乏持久记忆机制,最先进的视频对象分割技术依赖于反应式干扰物过滤而非强制全局身份一致性。为解决这些限制,我们提出REMIND,一种用于从单目RGB图像中对通用室内物体进行长期多目标重新识别的在线跟踪器,无需相机姿态和深度信息。受视觉认知中人类依赖累积外观熟悉度和空间上下文而非明确自我定位的证据启发,REMIND将冻结的DINOv3特征与双库多原型外观记忆、部分和背景级描述符、利用空间共现的邻居上下文推理模块以及具有模糊感知保障的联合匈牙利分配相结合。在一个具有受控重访和密集同类别杂波的专用室内数据集上,REMIND达到了90.35%的IDF1,比最先进的视频对象分割基线高出近20个点,比强大的检测跟踪基线高出36个以上。在ScanNet++上,除了一种设置(所有场景的端到端检测)外,它在每个设置中都获得了最高的IDF1,在该设置中检测跟踪基线略领先,但REMIND仍能更准确地关联和恢复身份;它还完成了每个场景,而视频对象分割基线在YOLO检测下66.9%的场景中耗尽了GPU内存。完整的系统、评估框架和数据集已公开发布。

英文摘要

Mobile robots operating indoors must re-identify previously observed objects after long temporal gaps, significant viewpoint changes, and severe illumination variations. This remains a challenging problem: multi-object tracking methods are optimized for short-term association of pedestrians and vehicles at video rates, person and vehicle re-identification approaches lack persistent memory mechanisms, and state-of-the-art video object segmentation techniques rely on reactive distractor filtering rather than enforcing global identity consistency. To address these limitations, we present REMIND, an online tracker designed for long-term multi-object re-identification of generic indoor objects from monocular RGB imagery, requiring neither camera pose nor depth. Motivated by evidence from visual cognition that humans rely on accumulated appearance familiarity and spatial context rather than explicit self-localization, REMIND combines frozen DINOv3 features with a dual-bank multi-prototype appearance memory, part- and background-level descriptors, a neighbour-context reasoning module exploiting spatial co-occurrence, and joint Hungarian assignment with ambiguity-aware safeguards. On a purpose-built indoor dataset featuring controlled revisits and dense same-class clutter, REMIND reaches 90.35% IDF1, nearly 20 points above a state-of-the-art video object segmentation baseline and more than 36 above a strong tracking-by-detection baseline. On ScanNet++, it attains the highest IDF1 in every setting but one, end-to-end detection over all scenes, where the tracking-by-detection baseline is marginally ahead while REMIND still associates and recovers identities more accurately; it also completes every scene, whereas the video object segmentation baseline exhausts GPU memory on 66.9% under YOLO detections. The complete system, evaluation framework, and dataset are publicly released.

发表机构

  • Centre for Automation and Robotics C.A.R. (UPM-CSIC)(自动化与机器人技术中心C.A.R.(马德里理工大学 - 西班牙国家研究委员会))
  • Universidad Politécnica de Madrid(马德里理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑