arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CHARM:基于大语言模型的多模态讽刺检测中的电荷校准与声学救援

CHARM: Charge Calibration and Acoustic Rescue for LLM-based Multimodal Sarcasm Detection

Qiyang Sun, Yi Chang, Yupei Li, Xi Shao, Zixing Zhang, Björn W. Schuller

arXiv 2607.11102首次发表:更新:

发表机构

GLAM – the Group on Language, Audio, & Music, Imperial College London; College of Telecommunications and Information Engineering, Nanjing University of Posts and Telecommunications; College of Computer Science and Electronic Engineering, Hunan University; Shenzhen Research Institute, Hunan University; CHI – Chair of Health Informatics, TUM University Hospital; relAI – the Konrad Zuse School of Excellence in Reliable AI; MDSI – Munich Data Science Institute; MCML – Munich Center for Machine Learning(伦敦帝国理工学院语言、音频和音乐小组; 南京邮电大学通信与信息工程学院; 湖南大学计算机科学与电子工程学院; 湖南大学深圳研究院; 慕尼黑工业大学医院健康信息学主席; 康拉德·楚泽可靠人工智能卓越学院; 慕尼黑数据科学研究所; 慕尼黑机器学习中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对大语言模型在讽刺检测中过度预测积极类别及韵律线索利用不足的问题,提出CHARM框架,含双向电荷校准和声学后期融合救援两个模块,无需微调主干,提升检测性能,揭示跨文化韵律解耦,产生可解释的跨语言多模态检测器。

AI 中文摘要

讽刺检测是情感计算中的一项基本任务。然而,零样本指令微调的大语言模型(LLMs)在整个能力范围内系统性地过度预测积极(讽刺)类别,而人类依赖的韵律线索未得到充分利用且跨语言转移不均衡。我们引入了CHARM(用于多模态讽刺检测的电荷校准与声学救援),这是一个无需训练的框架,它结合了两个模块。双向电荷校准(BiCAL)沿着带电提示的对称轴引导大语言模型做出相反的讽刺和字面判断;诱导的方向偏差通过构造相互抵消,简单聚合可恢复无偏的语用信号。声学后期融合救援(ALFR)然后通过一个浅层分类器将校准后的投票与韵律描述符和大语言模型生成的听觉感知探针融合,积极降低饱和文本投票的权重以支持声学证据。在不微调任何主干的情况下,BiCAL在MUStARD上实现了报告的最高零样本纯文本Macro-F1为0.787,而ALFR在CMMA上使弱主干的Macro-F1提高了多达+0.382。斯托弗荟萃分析证实了在MUStARD和CMMA上的统计显著性。我们的分析还发现了跨文化韵律解耦:低级声学特征无法跨语言转移,而高级感知抽象则保持稳健。这些组件共同产生了一个可解释的跨语言多模态检测器。

英文摘要

Sarcasm detection, the identification of discrepancies between literal and intended meaning, is a fundamental task in affective computing. However, zero-shot instruction-tuned Large Language Models (LLMs) systematically over-predict the positive (sarcastic) class across the entire capability spectrum, while the prosodic cues humans rely on remain underexploited and transfer unevenly across languages. We introduce CHARM (Charge Calibration and Acoustic Rescue for Multimodal Sarcasm Detection), a training-free framework that couples two modules. Bidirectional Charge Calibration (BiCAL) steers the LLM toward opposing sarcastic and literal verdicts along a symmetric axis of charged prompts; the induced directional biases cancel by construction, and a simple aggregation recovers an unbiased pragmatic signal. Acoustic Late-Fusion Rescue (ALFR) then fuses the calibrated votes with prosodic descriptors and LLM-generated auditory-perception probes through a shallow classifier, actively down-weighting saturated text votes in favour of acoustic evidence. Without fine-tuning any backbone, BiCAL attains the highest reported zero-shot text-only Macro-F1 of 0.787 on MUStARD, while ALFR lifts weak backbones by up to +0.382 Macro-F1 on CMMA. A Stouffer meta-analysis confirms statistical significance on MUStARD and CMMA (Z = 13.89 and Z = 34.64, respectively; p < 10^-43). Our analysis further uncovers a cross-cultural prosodic decoupling: low-level acoustics fail to transfer across languages, whereas high-level perceptual abstractions remain robust. Together, these components yield an explainable, cross-lingual multimodal detector.

Commentsunder review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑