发表机构
Sichuan University; Beijing Institute of Technology(四川大学; 北京理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉否定理解中现有模型与指标的失效问题,提出SNUS任务及CESG评估指标,构建高保真数据集,为安全关键场景提供可靠基准。
AI 中文摘要
真正的机器智能需要超越被动的像素登记,通过视觉否定理解掌握对缺失信息的自上而下的功能性推理。然而,无约束的视觉否定范式仍然过于开放,且普遍的肯定偏差导致现有的多模态大语言模型(MLLMs)和评估指标在否定语义下均失效。为系统性地解决这些相互交织的挑战,我们首先将否定推理的边界锚定在特定的认知目标内。具体而言,通过将安全作为高度实用且关键的认知维度,我们定义了安全认知下的场景否定理解(SNUS)任务。在该框架下,我们构建了一个高保真度的否定标题数据集,映射局部危险的密集断言。同时,我们提出了认知期望场景图(CESG)分数,这是一种基于结构、极性感知的评估指标。大量实验表明,尽管当前模型在此任务上表现困难,传统指标在语义反转下完全崩溃。相反,我们的框架为SNUS提供了坚实的基准,为推进风险感知的情境理解和反事实认知提供了严谨的基础。
英文摘要
True machine intelligence requires transcending passive pixel registration to master top-down functional reasoning over absent information via visual negation understanding. However, unconstrained visual negation paradigms remain overly open-ended, and pervasive affirmation bias causes both existing Multi-Modal Large Language Models (MLLMs) and evaluation metrics to fail under negative semantics. To solve these intertwined challenges systematically, we first anchor the boundaries of negation reasoning within specific cognitive goals. Specifically, by focusing on safety as a highly pragmatic and critical cognitive dimension, we define the task of \textbf{S}cene \textbf{N}egation \textbf{U}nderstanding under \textbf{S}afety Cognition (\textbf{SNUS}). Under this framework, we construct a high-fidelity negative caption dataset mapping dense assertions of localized hazards. Concurrently, we propose the Cognitive Expected Scene Graph (CESG) Score, a structure-grounded, polarity-aware evaluation metric. Extensive experiments demonstrate that while current models struggle on the task, traditional metrics completely collapse under semantic reversals. Conversely, our framework delivers a solid benchmark for SNUS, providing a rigorous foundation to advance risk-aware situational comprehension and counterfactual cognition.