SOS!:用于无模型分割的精简式物体条件Transformer
SOS! : A Streamlined Object-Conditional Transformer for Model-free Segmentation
浏览论文内容
中文总结 AI 辅助
针对基础分割模型难以关联掩码与目标物体的问题,提出SOS框架,以单次前向传播统一掩码生成与目标识别,在无模型未见过物体分割上达当前最佳性能
中文摘要 AI 辅助
基础分割模型擅长生成高质量的类别无关掩码,但难以将这些候选结果与特定目标物体关联,这种语义鸿沟严重阻碍了它们在机器人操纵等下游应用中的部署,此类应用要求对未见过的物体进行精确分割。现有方法试图通过依赖详尽的3D物体模型先验来解决该问题,但这必然会带来过高的计算开销和复杂的多阶段流程。为应对这些局限,我们提出SOS(用于无模型分割的精简式物体条件Transformer)。SOS完全消除了对3D模型的依赖,每个目标物体仅需一张参考图像。我们框架的核心是一种新型的物体条件Transformer,它学习以身份为锚点的查询,将掩码生成和目标识别统一为单次前向传播。这种精简设计大幅提升了结构和计算效率。在多个基准上的广泛评估表明,SOS在无模型未见过物体分割领域达到了新的当前最佳水平,兼具精确性和高效率。项目页面和代码可在该https URL获取。
英文摘要
Foundation segmentation models excel at generating high-quality, class-agnostic masks, but they struggle to associate these proposals with specific target objects. This semantic gap severely hinders their deployment in downstream applications like robotic manipulation, which demand precise unseen objects segmentation. Existing approaches attempt to resolve this by relying on exhaustive 3D object model priors, inherently introducing prohibitive computational overhead and complex, multi-stage pipelines. To address these limitations, we propose SOS (Streamlined Object-conditional Transformer for model-free Segmentation). SOS completely eliminates the reliance on 3D models, requiring only a single reference image per target object. Central to our framework is a novel Object-Conditional Transformer that learns identity-anchored queries, unifying mask generation and target identification into a single feed-forward pass. This streamlined design drastically improves both structural and computational efficiency. Extensive evaluations across multiple benchmarks demonstrate that SOS establishes a new state-of-the-art for model-free unseen objects segmentation, delivering accurate and high-efficiency performance. The project page and code are available at https://sos-seg.github.io/.
发表机构
- Technical University of Munich(慕尼黑工业大学)
- Siemens AG(西门子公司)
- Munich Center for Machine Learning(慕尼黑机器学习中心)
- ROBOX
机构由 AI 辅助整理,请以论文原文为准。