CLASH:口语讽刺检测中词汇与韵律依赖的反事实审计
CLASH: Counterfactual Auditing of Lexical and Prosodic Reliance in Spoken Sarcasm Detection
浏览论文内容
中文总结 AI 辅助
CLASH提出双语反事实诊断框架,审计口语讽刺检测中词汇与韵律的依赖,揭示模型主要依赖词汇线索,区分声学敏感性与讽刺判别。
中文摘要 AI 辅助
口语讽刺检测器可能利用词汇内容、韵律或其交互作用,然而传统评估无法揭示哪些线索驱动其预测。我们提出CLASH(受控词汇-声学分离框架),一个双语反事实诊断框架,在原始、词汇保持、韵律保持和近似中性化条件下评估每个话语。我们在CMMA和MUStARD上评估了手工声学特征系统、自监督学习(SSL)探针和大型音频语言模型(LALMs)。对于仅目标Qwen3-Omni,在时长平衡后,词汇保持语音相比韵律保持语音保留了0.135--0.148的AUROC优势,聚类自助区间高于零;替代词汇重合成保留了这一优势。声学干预改变了分数,但在评估条件下未持续改善判别能力或改变二元预测。上下文和交互估计因语料库而异。这些发现区分了声学敏感性与讽刺判别,同时揭示了时长、身份和转换效应。
英文摘要
Spoken sarcasm detectors may exploit lexical content, prosody, or their interaction, yet conventional evaluation cannot reveal which cues drive their predictions. We introduce CLASH (Controlled Lexical-Acoustic Separation Harness), a bilingual counterfactual diagnostic framework that evaluates each utterance under original, lexical-preserving, prosody-preserving, and approximately neutralised conditions. We evaluate handcrafted acoustic-feature systems, self-supervised learning (SSL) probes, and large audio language models (LALMs) on CMMA and MUStARD. For target-only Qwen3-Omni, lexical-preserving speech retains a 0.135--0.148 AUROC advantage over prosody-preserving speech after duration balancing, with cluster-bootstrap intervals above zero; alternative lexical resynthesis preserves this advantage. Acoustic interventions shift scores without consistently improving discrimination or changing binary predictions under the evaluated conditions. Context and interaction estimates vary across corpora. These findings distinguish acoustic sensitivity from sarcasm discrimination while exposing duration, identity, and transformation effects.
发表机构
- Imperial College London(伦敦帝国理工学院)
- New York University(纽约大学)
- Technical University of Munich(慕尼黑工业大学)
- Tencent Inc.(腾讯公司)
- Wuhan University(武汉大学)
- Nankai University(南开大学)
机构由 AI 辅助整理,请以论文原文为准。