发表机构
Huazhong University of Science and Technology; Beijing Normal University; Waseda University(华中科技大学; 北京师范大学; 早稻田大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究能否仅用RGB图像训练自动抠图模型。提出自监督框架SSMatte,将问题分解为语义锚定和细节抠图,通过特定损失生成提示并优化网络。该方法性能优,推动自动抠图进入完全无标注范式。
AI 中文摘要
高质量的alpha遮罩标注成本高昂,这为深度图像抠图造成了根本性的数据瓶颈。尽管先前工作试图使用诸如三通道图或掩码等更粗糙的标签来降低标注成本,但它们仍依赖昂贵的逐像素监督,限制了可扩展性和泛化能力。在本文中,我们进一步拓展边界并提出问题:能否仅使用RGB图像训练自动抠图模型,完全无需人工标注?我们通过展示SSMatte来回答这一问题,它是一个自监督框架,首次实现了与全监督自动抠图相当的性能。我们的关键见解是将问题分解为语义锚定和细节抠图。SSMatte首先通过基于广义瑞利商的新颖且训练高效的语义锚定损失传播类令牌种子,从冻结的自监督ViT特征生成语义抠图提示。然后该提示锚定一个细节抠图网络,通过基于定点的损失进行优化以强制alpha-RGB一致性。大量实验表明,SSMatte优于先前的弱监督方法,在人像基准测试中与全监督模型的性能相匹配,并在增加数据时展现出良好的可扩展性和泛化行为。我们的工作将自动抠图推向了全新的、完全无需标注的范式。代码将公开。
英文摘要
High-quality alpha mattes are notoriously expensive to annotate, creating a fundamental data bottleneck for deep image matting. While prior work attempts to reduce annotation cost using coarser labels like trimaps or masks, they remain reliant on costly per-pixel supervision, limiting scalability and generalization. In this work, we push the boundary further and ask: can we train an automatic matting model using only RGB images, with no manual annotation at all? We answer this by presenting SSMatte, a self-supervised framework that for the first time achieves performance on par with fully-supervised automatic matting. Our key insight is to decompose the problem into semantic anchoring and detail matting. SSMatte first generates a semantic matting prompt from frozen self-supervised ViT features by propagating class-token seeds via a novel, training-efficient semantic anchoring loss based on a generalized Rayleigh quotient. This prompt then anchors a detail matting network, which is optimized via a fixed-point-based loss that enforces alpha-RGB consistency. Extensive experiments show SSMatte outperforms prior weakly-supervised methods, matches the performance of fully-supervised models on portrait benchmarks, and demonstrates favorable scaling and generalization behaviors with additional data. Our work pushes automatic matting to an fresh, fully annotation-free paradigm. Code will be available.