发表机构
South China University of Technology; Shenzhen MSU-BIT University; Tsinghua University(华南理工大学; 深圳北理莫斯科大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉情感识别中仅凭表象证据不足的问题,提出事件锚定情感识别EGER及无需微调的AffectReveal框架,通过恢复并验证决定情感的事件,显著提升识别准确率。
AI 中文摘要
视觉情感识别通常假设预测所需的全部证据都包含在观察到的图像或视频中。然而,相同的可见反应可能传达不同的情感,这取决于输入之外的事件:例如,眼泪可能表示悲伤或喜悦。我们提出了事件锚定情感识别(EGER),其中情感识别需要恢复决定情感的事件。我们构建了EGER-Bench基准,包含10,052个视频和10,734张图像,涵盖11种情感、两个源域和四种视觉设置。一项由六名标注者参与的研究表明,事件上下文将人类识别准确率从33.96%提高到72.08%,证实了仅凭视觉证据往往是不够的。仅凭语义相关性并不能解决EGER:如果一个可能的事件其身份、焦点人物角色、关系或结果被误解,那么该事件可能暗示错误的情感。因此,我们提出了AffectReveal,一个无需微调的框架,它首先构建并独立验证基于证据的替代方案,涵盖这些情感关键因素。然后,它通过双向原子证据支持,将恢复的事件与面部遮挡的媒体内事实进行交叉核对,同时保留原始未遮挡输入用于最终预测。在三个下游模型和四种输入设置中,AffectReveal平均UAR提升了5.26至10.53个百分点。对于三个可微调模型,它还能使未微调模型在所有12项准确率比较中优于其微调的纯视觉对应模型,而无需更新下游参数。
英文摘要
Visual emotion recognition commonly assumes that all evidence required for prediction is contained in the observed image or video. Yet the same visible reaction can convey different emotions depending on events beyond the input: tears, for example, may indicate grief or joy. We formulate Event-Grounded Emotion Recognition (EGER), where emotion recognition requires recovering the affect-determining event. We construct EGER-Bench, comprising 10,052 videos and 10,734 images across 11 emotions, two source domains, and four visual settings. A study with six annotators shows that event context raises human recognition accuracy from 33.96% to 72.08%, confirming that visual evidence alone is often insufficient. Semantic relevance alone does not solve EGER: a plausible event may imply the wrong emotion if its identity, focal-person role, relationship, or outcome is misinterpreted. We therefore propose AffectReveal, a tuning-free framework that first constructs and independently verifies evidence-grounded alternatives over these affect-critical factors. It then cross-checks the recovered event against face-masked in-media facts through bidirectional atomic evidence support, while retaining the original unmasked input for final prediction. Across three downstream models and four input settings, AffectReveal yields average UAR gains of 5.26--10.53 points. For three fine-tunable models, it also enables untuned models to outperform their fine-tuned visual-only counterparts in all 12 accuracy comparisons, without updating downstream parameters.
Comments23 pages, 5 figures, 7 tables. Submitted to ICLR 2027