AI 中文总结
研究针对RGB-D语义分割中模态缺失问题,提出条件丢弃(ConD)持续训练范式,从预训练模型出发,通过特定模拟输入、冻结及训练复制编码器等操作,提升模型在缺失模态下的鲁棒性,完整模态时也有增益。
AI 中文摘要
RGB-D语义分割已取得显著进展,但大多数模型假定RGB和深度总是可用的。在实际中,监控传感器的故障或遮挡常导致一种模态缺失。尽管仅RGB或深度就可能包含足够线索,但仅在全模态输入上训练的模型在一种模态缺失时无法利用剩余模态,导致性能严重下降。我们用一种简单的持续训练范式——条件丢弃(ConD)来解决这个问题,它在保持全模态准确性的同时减轻性能下降。从预训练的RGB-D模型开始,ConD添加第二阶段,随机模拟完整、RGB缺失和深度缺失的输入,冻结原始编码器,并用零初始化特征注入训练复制的编码器。在NYU-Depth V2和SUN RGB-D上的实验表明,ConD提高了在缺失模态下的鲁棒性,甚至在模态完整时也有轻微提升。我们的代码将在被接受后公开。
英文摘要
RGB-D semantic segmentation has achieved remarkable progress, yet most models assume that RGB and depth are always available. In practice, failures or occlusions of surveillance sensors often remove one modality. Although RGB or depth alone can contain sufficient cues, models trained only on full-modality inputs fail to exploit the remaining modality once one is missing, causing severe degradation. We tackle this issue with a simple continued-training paradigm, \emph{Condition Dropout (ConD)}, which mitigates degradation while preserving full-modality accuracy. Starting from a pretrained RGB-D model, ConD adds a second stage that randomly simulates complete, RGB-missing, and depth-missing inputs, freezes the original encoders, and trains copied encoders with zero-initialized feature injection. Experiments on NYU-Depth V2 and SUN RGB-D show that ConD improves robustness under missing modalities and even yields slight gains when modalities are complete. Our code will be made publicly available upon acceptance.