arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04939cs.CL

帧间解读:社交媒体视频中隐含与非字面意义的理解

Reading Between the Frames: Interpreting Implicit and Non-literal Meaning in Social Media Videos

Yang Wang, Yanan Ma, Yiqi Liu, Zi Yan Chang, Chi-Li Chen, Chia-Yi Hsiao, Tyler Loakman, Aline Villavicencio, Chenghao Xiao, Chenghua Lin

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出DrivelHub+基准,评估视频-语言模型推断社交媒体视频隐含意义的能力,从解释和表征两方面开展实验,衡量模型多模态感知与语用理解的差距。

中文摘要 AI 辅助

社交媒体视频常传递超出可见动作、字幕或语音的含义,一段普通片段仅通过多模态线索与文化语境的互动就可能变得幽默、讽刺或具有讽刺意味,这类内容对视频-语言模型而言是极具挑战性的测试案例。本文提出DrivelHub+,这是一个用于评估模型能否推断社交媒体视频隐含、非线性且修辞分层意义的基准,这类视频表面看似无意义,但传递着刻意的语用意义。DrivelHub+包含从社交媒体收集的1000个视频,每个视频都附有人类编写的隐含叙事解释。与聚焦识别或描述的传统视频理解任务不同,我们提出的基准针对的是情境多模态推理。我们从两个角度评估当前的视频-语言模型:一是解释任务,要求模型用自然语言解释对视频的语用理解;二是表征任务,我们将推理即检索的方法适配到该任务,以测试模型表征在视频到文本和文本到视频检索中是否能将视频与其对应的隐含叙事对齐。我们的基准提供了一个诊断场景,用于衡量多模态感知与语用理解之间的差距,探究当前模型能否超越对可见内容的描述,进而推断视频的隐含意义。

英文摘要

Social media videos often communicate meanings that go beyond their visible actions, captions, or speech. A mundane clip may become humorous, ironic, or satire only through the interaction of multimodal cues and cultural context, making such content a difficult test case for video-language models. In this paper, we introduce \textit{DrivelHub+}, a benchmark for evaluating whether models can infer the implicit, non-linear, and rhetorically layered meanings of social media videos that appear nonsensical on the surface but convey deliberate pragmatic meanings. DrivelHub+ consists of 1,000 videos collected from social media, each annotated with a human-written implicit narrative explanation. Unlike conventional video understanding tasks focused on recognition or description, we present a benchmark that targets contextual multimodal reasoning. We evaluate current video-language models from two perspectives: explanation, where models must explain the pragmatic comprehension of a video in natural language; and representation, where we adapt reasoning-as-retrieval to test whether model representations align videos with their corresponding implicit narratives in both video-to-text and text-to-video retrieval. Our benchmark provides a diagnostic setting for measuring the gap between multimodal perception and pragmatic comprehension, asking whether current models can move beyond describing what is shown to inferring what is meant.

发表机构

  • University of Manchester(曼彻斯特大学)
  • Durham University(杜伦大学)
  • University of Sheffield(谢菲尔德大学)
  • University of Exeter(埃克塞特大学)

机构由 AI 辅助整理,请以论文原文为准。

↑