情感的明暗对比:基于评价理论的对比情感基准测试集
Chiaroscuro for Emotions: A Contrastive Emotion Benchmark Grounded in Appraisal Theory
浏览论文内容
中文总结 AI 辅助
该研究提出基于评价理论的对比情感基准测试集CHIARO,经测试前沿大语言模型和现有情感分类器表现不佳,结合现有语料库训练后可提升情感识别性能。
中文摘要 AI 辅助
情感识别基准测试集通常为每段文本预测一种情感,却忽略了现实中常见的场景:同一事件会让不同的人产生截然相反的情感。例如,一个孩子兴奋地踢前面的座椅,而前排乘客则会因此生气。我们提出CHIARO,这是一个包含1000个人工标注句子的对比情感推理基准测试集,基于评价理论构建。每个场景描述一个因果触发因素,会让一个人产生积极情感,另一个人产生消极情感,这些场景来自一个十类分类体系。我们对七个前沿大语言模型(LLM)和四个现成的情感分类器进行了基准测试。表现最强的大语言模型达到了67.3的宏观F1值,远低于人类的一致性水平,而现有的情感分类器得分接近随机水平。除了用于评估外,CHIARO还可作为训练信号:当与现有的情感语料库结合时,得到的下游分类器在CHIARO自身及十个外部情感基准测试集中的六个上均有所提升,这表明我们的数据集可作为情感识别的补充信号。
英文摘要
Emotion recognition benchmarks often predict one emotion per text, missing many real-world scenarios where two people arrive at opposing emotions from a single shared event. For example, a child kicks the seat in front of her in excitement while the passenger ahead grows angry. We introduce CHIARO, a 1,000 human-annotated sentence benchmark for contrastive emotion inference grounded in appraisal theory. Each scene describes one causal trigger eliciting a positive emotion in one person and a negative emotion in the other, drawn from a ten-class taxonomy. We benchmark seven frontier LLMs and four off-the-shelf emotion classifiers. The strongest LLM reaches 67.3 macro-F1, well below human agreement, while existing emotion classifiers score near chance. Beyond evaluation, CHIARO also serves as a training signal. When combined with an existing emotion corpus, the resulting downstream classifier improves on CHIARO itself and on six of ten external emotion benchmarks, which positions our dataset as a complementary signal for emotion recognition.
发表机构
- University of Cincinnati(辛辛那提大学)
机构由 AI 辅助整理,请以论文原文为准。