发表机构
Sichuan University; Beijing Institute of Technology(四川大学; 北京理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对安全关键场景中视觉负事件理解难题,提出基于反事实重构与对比解码(CRCD)的负向描述框架,通过双分支重构和多条件表征学习缓解偏差,实现高效语义缺失推断并建立高性能基线。
AI 中文摘要
传统的场景理解聚焦于图像中客观存在的肯定性信息。然而,在安全关键领域中,理解本应存在但实际上缺失的关键信息对于风险缓解至关重要。为弥合这一差距,我们以安全作为认知约束,聚焦于视觉场景负向描述任务。核心挑战在于将物理上的缺失转化为语义上的负事件。现有的视觉-语言模型(VLMs)在此过程中表现不佳,因为肯定性偏差抑制了负向推理,而有限的心理填充能力和表征偏差进一步阻碍了对缺失信息的推断。为应对这些挑战,我们提出了一种基于反事实重构与对比解码(CRCD)的负向描述框架。受人类认知启发,CRCD将任务重构为反事实潜在变化描述,以绕过肯定性偏差。它将合成的安全预期与现实进行对比,以识别语义缺失。为解决有限的心理填充问题,我们设计了一种双分支反事实重构架构。其中,异质补全分支恢复缺损物体,而功能关联分支推断完全缺失的安全物体。同时,集成了多条件表征学习机制,通过将通用特征投影到预定义的安全准则子空间来缓解表征偏差,从而在更多维度上捕获信息。通过解码重构场景原型与原始输入之间的特征级语义残差,CRCD限定了非存在搜索空间,并激活了解码器的负向逻辑。大量实验验证了CRCD的有效性,为这一开创性任务建立了高性能基线。
英文摘要
Traditional scene understanding focuses on affirmative information objectively present in images. However, in safety-critical domains, comprehending key information that should exist but is actually absent is vital for risk mitigation. To bridge this gap, we focus on visual scene negative captioning with safety as the cognitive constraint. The core challenge is to convert physical absence into semantic negative events. Existing vision-language models (VLMs) struggle with this process because affirmation bias suppresses negative reasoning, while limited mental filling capability and representation bias further hinder the inference of absent information. To address these challenges, we propose a negative captioning framework based on counterfactual reconstruction and contrastive decoding (CRCD). Inspired by human cognition, CRCD reformulates the task as counterfactual latent change captioning to bypass affirmation bias. It contrasts a synthesized safe expectation with reality to identify semantic omissions. To address limited mental filling, we design a dual-branch counterfactual reconstruction architecture. The amodal completion branch restores defective objects, while the functional association branch infers completely absent safety objects. Concurrently, a multi-condition representation learning mechanism is integrated to mitigate representation bias by projecting universal features onto predefined safety criteria subspaces, thereby capturing information across more dimensions. By decoding feature-level semantic residuals between the reconstructed scene prototype and raw input, CRCD bounds the non-existence search space and activates the decoder's negative logic. Extensive experiments validate the effectiveness of CRCD, establishing a high-performance baseline for this pioneering task.