发表机构
Duke University; EPFL; National University of Singapore; Tsinghua University; Harvard University(杜克大学; 瑞士洛桑联邦理工学院; 新加坡国立大学; 清华大学; 哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
UnfoldCRF将掩码细化建模为条件随机场推理,引入图像条件潜在区域和空状态,通过展开平均场更新实现结构化细化,在COD10K及多数据集上显著超越黑盒对照组,并具有良好泛化性。
AI 中文摘要
学习型掩码细化器能够提升分割精度,但很难判断改进中有多少来自显式结构而非额外容量,以及在掩码生成器或其误差分布发生变化时该改进是否仍然成立。UnfoldCRF将细化视为在像素标签和潜在区域变量上的条件随机场中的推理。其能量包含修正的一元项、学习到的局部成对交互以及图像条件的潜在区域一致性,并设有一个空状态,使标签一致性较弱的区域能够退出一致性项;推理在该单一能量上展开阻尼平均场更新。为隔离结构的影响,我们与读取相同输入、获得相同参数预算、阶段数和监督的循环黑盒细化器进行比较。在COD10K上,UnfoldCRF比最强匹配对照组高出1.0个$F^\omega_\beta$点,改善了全部四项COD指标,并将变差图像比例从11.7%降至8.5%。在涵盖五个数据集和多种掩码来源的一次训练协议下,260万参数变体平均获得4.2个$\Delta$IoU增益,而对照组为2.0;基于冻结DINOv2特征的变体以约七分之一的驻留参数匹配最强基础模型细化器,同时领先于其自身对照组。在训练中从未见过的掩码生成器上,增益为2.0个$F^\omega_\beta$点,对照组为0.6。将单个消息置零可显示修正来源:成对消息主要修正边界,区域消息主要修正非边界错误。代码和辅助材料将公开发布。
英文摘要
Learned mask refiners improve segmentation accuracy, but it is hard to tell how much of the improvement comes from explicit structure rather than from extra capacity, and whether it holds up when the mask generator or its error distribution changes. UnfoldCRF treats refinement as inference in a conditional random field over pixel labels and latent region variables. Its energy has a corrected unary term, learned local pairwise interactions, and image-conditioned latent-region consistency, with a null state that lets a region with weak label agreement withdraw from the consistency term; inference unrolls damped mean-field updates on this one energy. To isolate the effect of structure, we compare against recurrent black-box refiners that read the same inputs and receive the same parameter budget, stage count, and supervision. On COD10K, UnfoldCRF beats the strongest matched control by 1.0 $F^ω_β$ point, improves all four COD metrics, and lowers the fraction of images made worse from 11.7\% to 8.5\%. Under a train-once protocol over five datasets and several mask sources, the 2.6M-parameter variant gains 4.2 mean $Δ$IoU against 2.0 for its control, and a variant built on frozen DINOv2 features matches the strongest foundation-model refiner with about a seventh of its resident parameters while staying ahead of its own control. On mask generators never seen in training, the gain is 2.0 $F^ω_β$ points against 0.6 for the control. Zeroing individual messages shows where the corrections come from: the pairwise messages mostly fix boundaries, the region messages mostly fix non-boundary errors. Code and supporting materials will be publicly released.
Comments16 pages