发表机构
Duke University; Carnegie Mellon University; University of Alabama at Birmingham(杜克大学; 卡内基梅隆大学; 阿拉巴马大学伯明翰分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出SphereTrust,利用冻结的DINOv2超球面特征对SAM候选伪掩码进行快速排序,并在多个分割池上超越基线,同时支持训练提升性能。
AI 中文摘要
诸如SAM等基础分割器会为未标注图像返回多个看似合理的掩码,而基于错误掩码训练的学生模型会继承其错误。在候选掩码中进行选择意味着要么查询另一个大型模型,要么为标注掩码拟合一个质量头。我们证明,可以通过候选掩码对冻结的自监督骨干网络特征的影响来判断其优劣。归一化的DINOv2补丁特征位于超球面上,而一个候选掩码将该球面一分为二。基于这一解读,我们提出了SphereTrust,它根据分割的三个属性对每个候选进行评分:两侧之间的角度对比度、前景外观模式的覆盖度以及与图像边框的接触,分别对应掩码失败的三种常见方式,并仅从冻结特征出发,以每张图像0.55秒的速度对候选池进行排序。在涵盖伪装、显著和二分分割以及低光伪装的八个SAM和SAM3候选池上,SphereTrust在六个池中的平均选定Dice系数上超过最强评估外部基线1.7至9.3个百分点。这些比较包括已发表的选取规则以及DSS和UCOD-MKD的明确标注改编版本。在两个提示伪装池上,其平均选定Dice系数与候选派生的DSS改编版本相差在0.1个百分点以内,且灾难性错误率更低。哪种线索携带信号取决于候选池。同一个球面也支持训练。领先的候选作为候选集进入,其分数作为先验,原型对它们重新排序,交叉拟合的第二轮完成标签,在三个MLLM锚点池上,与固定标签训练相比,加权F值提高了4.5、2.3和5.5个百分点,学生在十九个测试集上与已发表的无监督方法具有竞争力。
英文摘要
Foundation segmenters such as SAM return several plausible masks for an unlabeled image, and a student trained on the wrong one inherits its errors. Choosing among them means querying a second large model or fitting a quality head to annotated masks. We show that a candidate can be judged by what it does to a frozen self-supervised backbone's features. Normalized DINOv2 patch features lie on a hypersphere, and a candidate mask splits that sphere in two. Based on this reading, we introduce SphereTrust, which scores each candidate by three properties of the split, the angular contrast between the two sides, the coverage of the foreground's appearance modes, and contact with the image frame, one for each of three common ways a mask fails, and ranks a pool in 0.55 s per image from the frozen features alone. On eight SAM and SAM3 candidate pools spanning camouflaged, salient, and dichotomous segmentation and camouflage under low light, SphereTrust exceeds the strongest evaluated external baseline on six pools by 1.7 to 9.3 percentage points in mean selected Dice. These comparisons include published selection rules and explicitly labeled adaptations of DSS and UCOD-MKD. On the two prompted camouflage pools, its mean selected Dice is within 0.1 percentage points of the candidate-derived DSS adaptation, with a lower catastrophic-error rate. Which cue carries the signal depends on the candidate pool. The same sphere also supports training. The leading candidates enter as a candidate set with their scores as priors, prototypes reorder them, and a cross-fitted second round completes the labels, raising weighted F by 4.5, 2.3, and 5.5 points over fixed-label training on the three MLLM anchor pools, with students competitive with published unsupervised methods on nineteen test sets.
Comments29 pages, 11 figures, 19 tables