arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30204cs.CL

当模型“听到”它们所期待的:诊断多模态讽刺检测中的韵律启发式算法

When Models Hear What They Expect: Diagnosing Prosodic Heuristics in Multimodal Sarcasm Detection

Yongjian Chen, Pengfei Wei, Yiqun Sun, Zhu Li, Lawrence B. Hsieh

首次发表
浏览论文内容

中文总结 AI 辅助

本文以讽刺检测为切入点,评估Qwen2.5-Omni等多模态大语言模型,发现模型依赖音高升高、停顿不规则的韵律刻板印象,操控该维度可使假阳性率达60%,且该效应可跨模型复现。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)可联合处理语音与文本,但关于它们是否利用韵律线索进行语用推理,还是依赖表层声学模式,目前鲜有系统研究。本文以讽刺检测为切入点,在分解词汇内容、语音语义与韵律结构贡献的五种模态条件下,评估Qwen2.5-Omni与Qwen3-Omni在普通话和英语上的表现。加入音频会系统性增加假阳性,却未提升真阳性检测。声学错误诊断显示,模型错误集中于表达性韵律的共享刻板印象,即音高升高与停顿不规则,这与两种语言中标记讽刺的实际线索不符。仅对这两个维度进行针对性操控,可因果性证实该启发式算法,假阳性率最高达60%。将相同操控模板应用于Gemini 3 Flash Preview,无需修改即可复现该效应,表明该刻板印象不仅存在于Qwen Omni系列,并非单一模型架构所致。

英文摘要

Multimodal Large Language Models (MLLMs) process speech and text jointly, yet whether they exploit prosodic cues for pragmatic inference or rely on surface acoustic patterns has received little systematic investigation. We address this through sarcasm detection, evaluating Qwen2.5-Omni and Qwen3-Omni on Mandarin Chinese and English under five modality conditions that decompose the contributions of lexical content, vocal semantics, and prosodic structure. Adding audio systematically inflates false positives without improving true positive detection. Acoustic error diagnosis reveals that model errors cluster on a shared stereotype of expressive prosody, namely elevated pitch and irregular pausing, that diverges from the actual cues marking sarcasm in both languages. Targeted manipulation of only these two dimensions causally confirms the heuristic, inducing false positive rates of up to 60%. Applying the same manipulation template to Gemini~3 Flash Preview without modification replicates the effect, suggesting that the stereotype extends beyond the Qwen Omni family rather than arising from a single model architecture.

发表机构

  • Magellan Technology Research Institute (MTRI)(麦哲伦技术研究所(MTRI))
  • Center for Language and Cognition, University of Groningen(格罗宁根大学语言与认知中心)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑