AI 中文总结
针对以物体为中心模型生成掩码的缺陷,提出无训练后处理方法SSR,借助冻结自监督视觉Transformer优化掩码,在多类基准测试中显著提升全像素调整兰德指数,为掩码优化提供可迁移信号。
AI 中文摘要
以物体为中心的模型常产生碎片化掩码、边界泄漏及错误区域合并问题。本文提出相似度偏移优化(Similarity-Shift Refinement,SSR),一种无需训练的后处理方法,借助冻结的自监督视觉Transformer改进以物体为中心的掩码。SSR测量自注意力值聚合前后的成对块相似度变化,保留正向增强关系并构建稀疏亲和图,该图在单次优化步骤中传播初始软槽分配,无需重新训练或修改模型。在自然图像、合成视频及真实视频基准测试中,SSR在全部24个评估的模型-数据集组合中提升了全像素调整兰德指数,平均提升8.5个百分点。 ablation实验表明,值空间相似度偏移优于查询、键空间变体及静态Transformer亲和性,但纹理密集场景可能导致视觉相似区域被过度分组。总体而言,SSR为无训练以物体为中心的掩码优化提供了简单且可迁移的信号。
英文摘要
Object-centric models often produce fragmented masks, boundary leakage, and incorrect region merging. We introduce Similarity-Shift Refinement (SSR), a training-free post-hoc method for improving object-centric masks with a frozen self-supervised Vision Transformer. SSR measures changes in pairwise patch similarity before and after self-attention value aggregation, retains positively strengthened relations, and constructs a sparse affinity graph. This graph propagates the initial soft slot assignments in a single refinement step, without retraining or modifying either model. Across natural-image, synthetic-video, and real-world-video benchmarks, SSR improves all-pixel Adjusted Rand Index in all 24 evaluated model-dataset combinations, with an average gain of 8.5 percentage points. Ablations show that value-space similarity shifts outperform query- and key-space variants as well as static Transformer affinities. However, texture-dense scenes may cause visually similar regions to be over-grouped. Overall, SSR provides a simple and transferable signal for training-free object-centric mask refinement.