arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于可观测语义-图像接口与分层生成器证据对齐的生成式语义分割

Generative Semantic Segmentation via an Observable Semantic-Image Interface and Hierarchical Generator Evidence Alignment

Weize Cai, Yongqi Dong, Zhida Shao, Zixin Fu

arXiv 2608.11537首次发表:更新:

AI 中文总结

提出Semantic Prism框架,通过可观测语义-图像接口与分层生成器证据对齐实现生成式语义分割,在Cityscapes等数据集上提升了分割精度与像素误差排名性能。

AI 中文摘要

生成式语义分割将结构化预测以图像形式呈现,但直接颜色解码易受颜色漂移和边界混合影响,而预测独立输出分布的潜在特征解码器可能会将渲染图像降为中间可视化结果。我们提出Semantic Prism,一种具有确定性推理的条件语义-图像生成与优化框架。经扩散蒸馏的单步生成器渲染语义RGB图像;渲染颜色与固定类别码本的逐像素距离定义了显式概率接口。分层生成器证据对齐对多级生成器特征进行空间对齐,并使用零初始化输出投影预测接口对数几率空间中的加性残差,保留图像定义的接口作为最终分布的参考。该接口与优化后的分布进一步支持Contextual Interface--Hierarchy Disagreement(C-IHD),这是一种固定读出机制,无需辅助预测器或额外前向传播即可对剩余像素误差进行排名。在包含500张图像的Cityscapes验证集上,Semantic Prism达到72.07%的平均交并比,比直接接口解码高出11.39个平均交并比点,预期校准误差为0.41%。对三个随机种子进行的匹配容量消融实验验证了联合对齐多级证据的益处。单独训练的模型在BDD100K上达到62.22%的平均交并比,而在Cityscapes上训练的模型在无目标域适配的情况下,冻结源域迁移至带对应关系的恶劣条件数据集(ACDC)时达到46.89%的平均交并比。在所有三个数据集上,C-IHD在相同分割预测的像素误差排名的精确率-召回率曲线下面积(AUPR)均优于最大软max概率;在ACDC上,其将AUPR从0.6580提升至0.7557。

英文摘要

Generative semantic segmentation exposes structured predictions as images, but direct color decoding is susceptible to color drift and boundary mixing, whereas latent-feature decoders that predict a separate output distribution may relegate the rendered image to an intermediate visualization. We present Semantic Prism, a conditional semantic-image generation-and-refinement framework with deterministic inference. A diffusion-distilled one-step generator renders a semantic RGB image; per-pixel distances from the rendered colors to a fixed class-color codebook define an explicit probabilistic interface. Hierarchical Generator Evidence Alignment spatially aligns multi-level generator features and uses a zero-initialized output projection to predict an additive residual in the interface logit space, retaining the image-defined interface as the reference for the final distribution. The interface and refined distributions further enable Contextual Interface--Hierarchy Disagreement (C-IHD), a fixed readout for ranking remaining pixel errors without an auxiliary predictor or additional forward pass. On the 500-image Cityscapes validation set, Semantic Prism achieves 72.07% mean intersection over union, 11.39 mIoU points above direct-interface decoding, with 0.41% expected calibration error. Matched-capacity ablations over three seeds support the benefit of jointly aligned multi-level evidence. A separately trained model attains 62.22% mIoU on BDD100K, while the Cityscapes-trained model reaches 46.89\% mIoU under source-frozen transfer to the Adverse Conditions Dataset with Correspondences, without target-domain adaptation. Across all three datasets, C-IHD consistently improves the area under the precision--recall curve for pixel-error ranking over maximum softmax probability on the same segmentation predictions; on ACDC, it raises AUPR from 0.6580 to 0.7557.

Comments15 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑