发表机构
University of Augsburg; University of the Bundeswehr Munich(奥格斯堡大学; 慕尼黑联邦国防军大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对细粒度情绪识别基准,本文发现现成视觉语言模型通过从logits读取答案而非生成答案,即可匹配或超越专用微调模型,表明基准衡量的是情绪诱发而非感知,且该效应在合成与真实数据上均存在。
AI 中文摘要
细粒度情绪识别支持治疗工具和社交机器人,但它需要面部数据,这引发了隐私和数据保护方面的担忧。EmoNet-Face-HQ通过生成肖像来应对这一问题,这些肖像由专家按照40类分类法进行评分,该分类法远比通常的六到八种基本情绪更为细致。在其附带的协议下,视觉语言模型(VLM)在该分类法上得分不佳,该基准因此得出结论:需要专用的微调模型,即Empathic-Insight-Face(EIF;Small/Large)。我们证明,当答案不是生成而是从logits中读取(每个类别一个二元查询)时,现成的VLM能够匹配或超越该微调模型。我们保留基准的图像、分类法和评分,仅改变答案的读取方式。专家在其测量最可靠的五个类别上的一致性为κ_w=0.468。在生成方式下,十一个开放权重VLM中没有哪个区间完全高于该锚点(κ_w=0.268-0.486)。在验证方式下,所有十一个模型都超过了该锚点,每个模型都显著更好,κ_w=0.507-0.586。其中三个还显著优于EIF,后者为κ_w=0.551(Small;Large为0.534)。收益来自分级概率,而非提出是/否问题:作为对照,将这些相同概率阈值化为是/否会消耗平均收益的142%,并将二值化降至生成性诱发之下,κ_w=0.254-0.423。在真实照片(FACES)上的复现较弱且结果不一:在通过有效性门槛的十个模型中,六个获得收益,三个为中性至正面,一个为负面,因此该效应不仅限于合成数据。
英文摘要
Fine-grained emotion recognition supports therapy tools and social robots, but it needs facial data, which raises privacy and data-protection concerns. EmoNet-Face-HQ answers that with generated portraits, expert-rated over a $40$-category taxonomy far finer than the usual six to eight basic emotions. Under the protocol it ships with, vision-language models (VLMs) score poorly on that taxonomy, and the benchmark concludes that a dedicated fine-tuned model is necessary: Empathic-Insight-Face (EIF; Small/Large). We show that off-the-shelf VLMs match or beat that fine-tuned model when the answer is not generated but read from the logits, as one binary query per category. We keep the benchmark's images, taxonomy and ratings, and change only how the answer is read. Experts agree at $κ_w = 0.468$ on the five categories they measure most reliably. Generatively, no interval among eleven open-weight VLMs lies entirely above that anchor ($κ_w=0.268$-$0.486$). Under verification all eleven clear it, each of them significantly better at $κ_w=0.507$-$0.586$. Three also significantly beat EIF sitting at $κ_w = 0.551$ (Small; $0.534$ Large). The gain comes from the graded probability and not from asking a yes/no question: as a control, thresholding those same probabilities to yes/no costs 142% of the average gains and drops binarization below generative elicitation to $κ_w=0.254$-$0.423$. A replication on real photographs (FACES) is weaker and mixed: of the ten models that pass a validity gate, six gain, three are neutral to positive and one is negative, so the effect is not confined to synthetic data.
CommentsPreprint. 19 pages, 6 figures