发表机构
SperidLabs(斯佩里实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ENEAS是一种可文本提示的统一方法,通过扩展SeC架构并结合语义验证层,实现实例精确跟踪与开放概念发现,可区分真实实例与相似伪影,适用于视频、无序数据集等场景。
AI 中文摘要
我们提出了ENEAS,一种统一的、可文本提示的实例跟踪与语义发现方法。可文本提示的分割模型,包括SAM 3等最新基础模型,仍存在时间幻觉、空间碎片化和语义分类错误问题:当目标离开视野时无法报告其缺失;在极端特写时仅分割局部纹理而非完整目标;优先考虑视觉特征而非本体现实,导致雕像、绘画或反射等视觉相似的伪影被分割为目标实体。ENEAS作为单一方法具备两种功能:对唯一实例进行精确跟踪和高质量分割,以及对文本查询命名的每个实例进行开放概念发现,由语义验证层完成解析。对于跟踪,我们扩展了几何鲁棒的SeC架构(此前仅适用于点交互),添加文本提示适配器并利用其时间记忆,使目标在消失时不会漂移到干扰项,即使充满整个视野也能保持完整。对于发现,验证层将高速视觉嵌入匹配与条件VLM优化相结合,仅对模糊候选调用语义推理,过滤出仅视觉模型无法区分的本体错误,同时保持低延迟。ENEAS专为3D重建设计(其中单个分类错误的干扰项会损坏资产),可实现视频、广泛库及时间或空间无序数据集的高质量语义跟踪与分割,具备区分真实实例与其替身(外观相似但并非同一的事物)的能力。代码和模型可在该https URL获取。
英文摘要
We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-promptable segmentation models, including the latest foundation models such as SAM 3, still suffer from temporal hallucinations, spatial fragmentation, and semantic misclassification: they fail to report target absence when an object leaves the field of view, segment local textures instead of the complete object during extreme close-ups, and prioritize visual features over ontological reality, so that visually similar artifacts such as statues, paintings, or reflections are segmented as target entities. ENEAS works two ways from a single method: precise tracking and high-quality segmentation of a unique instance, and open-concept discovery of every instance a text query names, resolved by a semantic verification layer. For tracking, we extend the geometrically robust SeC architecture, previously limited to point interactions, with a text-prompting adapter and leverage its temporal memory, so that the target is held through disappearance without drifting to distractors and kept whole even when it fills the entire view. For discovery, the verification layer combines high-speed visual embedding matching with conditional VLM refinement, invoking semantic reasoning only for ambiguous candidates, which filters out the ontological errors that visual-only models cannot distinguish while keeping latency low. Designed with 3D reconstruction in mind, where a single misclassified distractor corrupts the asset, ENEAS unlocks high-quality semantic tracking and segmentation of video, of broad libraries, and of collections of temporally or spatially unordered data, together with the discrimination to tell true instances from their doppelgangers: things that look alike but are not the same. The code and models are available at https://github.com/speridlabs/eneas
Comments19 pages, 5 figures, 6 tables. Code and models: https://github.com/speridlabs/eneas