arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29917cs.CV

FoundYou:个性化分割与检索的统一模型

FoundYou: A Unified Model for Personalized Segmentation and Retrieval

Gabriele Trivigno, Marcos Alfaro, Claudia Cuttano, Gabriele Berton, Luis Payá, Carlo Masone

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出FoundYou统一框架,基于SAM 2的实例级线索实现个性化分割与检索,在相关基准上取得显著性能提升,模型规模小、速度快,且在类别级检索任务中也达最优。

中文摘要 AI 辅助

个性化分割与个性化检索均旨在跨不同图像识别同一物理对象,前者在目标图像中定位该对象,后者检索该对象出现的图像。尽管二者共享实例级目标,但这两项任务大多各自发展,采用不同解决方案。本研究提出FoundYou,这是一个统一框架,其构建基于如下观察:以保留视频帧间对象身份为目标训练的Segment Anything 2(SAM 2)模型,固有地捕获了实例级线索。我们利用这一特性匹配独立图像间的对象,使分割与检索成为同一实例对齐过程的两个输出结果。这一统一视角解锁了传统基准之外的新能力,包括少样本个性化检索以及带灵活提示的可提示个性化分割。大量实验表明,FoundYou相比统一方法和任务特定方法均取得一致提升,包括在PerMIS数据集上提升18.4 mIoU,在ILIAS数据集上提升17.8 mAP。其性能随参考图像数量增加而提升,且对较弱提示保持鲁棒性。除个性化任务外,FoundYou在类别级检索基准上也取得了最优结果。值得注意的是,本方法将SAM 2-small模型完全冻结,仅添加590万可训练参数,形成一个总参数为5200万的模型,与唯一的现有统一解决方案相比,速度提升超过75倍,规模缩小20倍。代码可在该https URL获取。

英文摘要

Personalized segmentation and personalized retrieval both aim to identify the same physical object across different images. While the former localizes the object within a target image, the latter retrieves images where it appears. Despite this shared instance-level objective, the two tasks have largely evolved separately and are addressed with distinct solutions. In this work, we introduce FoundYou, a unified framework built on the observation that Segment Anything 2 (SAM 2), trained to preserve object identity across video frames, inherently captures instance-level cues. We leverage this property to match objects across independent images, enabling segmentation and retrieval to emerge as two outcomes of the same instance alignment process. This unified view unlocks new capabilities beyond traditional benchmarks, including few-shot personalized retrieval and promptable personalized segmentation with flexible prompts. Extensive experiments show consistent gains over unified and task-specific methods, including +18.4 mIoU on PerMIS and +17.8 mAP on ILIAS. Performance scales with additional references and remains robust to weaker prompts. Beyond personalization, FoundYou achieves state-of-the-art results on category-level retrieval benchmarks. Notably, our approach keeps the SAM 2-small model entirely frozen and adds only 5.9 M trainable parameters, yielding a 52 M-parameter model that is over 75x faster and 20x smaller than the only prior unified solution. Code is available at https://github.com/ga1i13o/FoundYou .

发表机构

  • Politecnico di Torino(都灵理工大学)
  • Miguel Hernández University of Elche(埃尔切米格尔·埃尔南德斯大学)
  • Valencian Graduate School and Research Network of Artificial Intelligence(瓦伦西亚人工智能研究生学院与研究网络)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑