arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习检测跨模态否定:潜在表示分析与基于注意力的解决方案

Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution

Ali AbuSaleh, Leon Hammerla, Alexander Mehler

arXiv 2607.17712首次发表:更新:

AI 中文总结

研究跨模态否定检测难题,提出新型跨模态注意力架构,分析发现文本与视觉否定的不对称性,结合自监督视频表示推进时间否定建模,为多模态系统学习语义对齐表示提供新方法。

AI 中文摘要

检测跨模态的否定等高级语义概念对当前多模态系统仍是挑战。本文将其视为基本表示学习问题,首次证明在标准视觉语言模型的潜在空间中,否定不构成线性或非线性可分的类别。预训练嵌入主要编码特定模态特征,缺乏可泛化的否定信号。为此提出新型跨模态注意力架构,明确建模模态间依赖关系,比单模态基线F1性能提升高达7.03%。分析揭示关键不对称性,文本否定常独立出现,视觉否定语义上依赖语言上下文,通过对3222个政治视频-文本对统计分析验证。结合此分析与自监督视频表示推进了时间否定建模,为多模态系统学习鲁棒、语义对齐表示提供新方法和见解。

英文摘要

Detecting high-level semantic concepts like negation across modalities remains a challenge for current multimodal systems. We analyze this as a fundamental representation learning problem, providing the first evidence that negation does not form a linearly or non-linearly separable class in the latent spaces of standard vision-language models (VLMs). We demonstrate that pretrained embeddings primarily encode modality-specific features, lacking a generalizable negation signal. To overcome this, we propose a novel cross-modal attention architecture that explicitly models inter-modal dependencies, achieving performance gains of up to +7.03% F1 over unimodal baselines. Our analysis reveals a key asymmetry: while textual negation often appears independently, visual negation is semantically dependent on linguistic context, a finding validated through our statistical analysis of 3,222 political video-text pairs automatically annotated via \textsc{Qwen2.5-VL}. By combining this analysis with self-supervised video representations (JEPA2), we advance the modeling of temporal negation. This work provides new methods and insights for learning robust, semantically-aligned representations in multimodal systems.

CommentsThis manuscript is an accepted version of the article (published at ICNLP2026). Published in IEEE Xplore, DOI:10.1109/ICNLP69856.2026.11527861 document: https://ieeexplore.ieee.org/abstract/document/11527861

Journal ref2026 8th International Conference on Natural Language Processing (ICNLP), Xi'an, China, 2026, pp. 613-622

DOI:10.1109/ICNLP69856.2026.11527861

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑