iSEE:通过自监督实现对象永久性
iSEE: Object Permanence Through Self-Supervision
浏览论文内容
中文总结 AI 辅助
iSEE通过自监督实现对象永久性,利用槽注意力、外观位置分离和重新出现训练,在无标签情况下提升遮挡跟踪性能。
中文摘要 AI 辅助
对象永久性,即在对象被遮挡时保持对其身份和位置的追踪,对于视频表示中的跟踪、预测和规划至关重要。实现这一目标的跟踪器从边界框、轨迹身份和可见性标签中学习。另一方面,自监督的以对象为中心的方法无需标签即可发现对象:通过槽注意力,它将视频表示为绑定到对象并跨帧跟踪它们的槽。然而,这些槽在遮挡下会丢失,使得所需的永久性无法实现。推理永久性是一个难题,因为它需要检测对象何时被遮挡,在对象重新出现时重新识别,并保持对象的隐藏位置连续,仅以重新出现作为唯一的学习线索。为了解决这个问题,我们提出了iSEE,一个新颖的框架,提供了上述所有三个要求,且无需任何标签。我们使用以下三个提出的组件构建了iSEE:(i)对象证据建模:槽的注意力与其自身过去相比,揭示了其对象何时被隐藏。(ii)外观-位置分离:两个槽流让外观被保留用于重新识别,而位置则持续变化。(iii)从重新出现中获得永久性:一个行走器跟踪隐藏对象的位置,仅基于对象重新出现的位置进行训练。在LA-CATER静态数据集上,iSEE在86%的遮挡后将重新出现的对象返回其自身的槽,而SlotContrast为32%,并在隐藏时将其定位在标签训练的SoTA RAM的4.1 mAP范围内。这两个流还允许下游规划,位置流作为世界模型的动作。项目页面:此https URL
英文摘要
Object permanence, keeping track of an object's identity and position while it is occluded, is central to video representations that track, predict and plan. Trackers that achieve it learn from boxes, track identities and visibility labels. On the other hand, self-supervised object-centric methods discover objects without labels: through slot attention, it represents a video as slots that bind to objects and follow them across frames. However, these slots are lost under occlusion, making the desired permanence impossible. Reasoning permanence is a hard problem because it requires to detect when an object becomes occluded, re-identify when object reappears, and keep the object's hidden position continuous, using reapperance as the only learning cue. To address this, we propose iSEE, a novel framework that offers all three aforementioned requirements, without any labels whatsoever. We built iSEE using the following three proposed components: (i) Object evidence modelling: a slot's attention, compared with its own past, reveals when its object is hidden. (ii) Appearance-position separation: two slot streams let the appearance be held for re-identification while the position keeps changing. (iii) Permanence from reappearance: a walker follows the hidden object's position, trained only on where the object reappears. On LA-CATER static, iSEE returns a reappearing object to its own slot after 86% of occlusions, against 32% for SlotContrast, and localises it while hidden within 4.1 mAP of the label-trained SoTA RAM. The two streams also allow downstream planning, with the position stream as the action of a world model. Project page: https://insait-institute.github.io/iSEE/
发表机构
- INSAIT, Sofia University “St. Kliment Ohridski”(索非亚大学“圣克莱门特·奥赫里德”INSAIT)
机构由 AI 辅助整理,请以论文原文为准。