arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LG-GER:基于多模态证据蒸馏的语言引导群体情绪识别

LG-GER: Language-Guided Group Emotion Recognition via Multimodal Evidence Distillation

Ahmed Shehab Khan, Zhiyuan Li, Yan Tong

arXiv 2608.23880首次发表:更新:

发表机构

University of South Carolina(南卡罗来纳大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出LG-GER框架,通过MLLM生成结构化证据并蒸馏至VLM骨干,无需检测器等即可实现高效群体情绪识别,在GroupEmoW和GAF 3.0数据集上取得了有竞争力或更优的结果。

AI 中文摘要

从单张图像推断人群的集体情绪状态,这一任务被称为群体情绪识别(Group Emotion Recognition, GER),需要整合面部、姿态、互动以及场景上下文等空间分布的线索。当前方法依赖检测器驱动的多流管道,这些方法仅使用图像级监督进行训练,缺乏对哪些区域重要或各区域贡献强度的引导。我们提出LG-GER,一种语言引导的蒸馏框架,该框架利用多模态大语言模型(Multimodal Large Language Model, MLLM)为训练图像生成密集的、空间定位的证据,即与情绪信号和置信度评分配对的边界框。该结构化证据通过四种互补损失被蒸馏到单个视觉语言模型(Vision-Language Model, VLM)骨干中:分类损失、区域-文本对齐损失、空间情绪损失以及空间置信度回归损失。在推理阶段,LG-GER不需要检测器、不需要MLLM,也不需要多流融合,使得GER能够在实时和资源受限的部署场景中应用。LG-GER已在两个基准GER数据集(GroupEmoW和GAF 3.0)上进行评估,与需要在推理阶段进行检测和多流处理的现有最先进方法相比,取得了具有竞争力或更优的结果。

英文摘要

Inferring the collective emotional state of a group of people from a single image, a task known as group emotion recognition (GER), requires integrating spatially distributed cues such as faces, poses, interactions, and scene context. Current methods rely on detector-driven multi-stream pipelines. These are trained with only image-level supervision that lacks guidance on which regions matter or how strongly each contributes. We propose LG-GER, a language-guided distillation framework that uses a multimodal large language model (MLLM) to generate dense, spatially grounded evidence, i.e., bounding boxes paired with emotion signals and confidence scores, for the training images. This structured evidence is distilled into a single vision-language model (VLM) backbone through four complementary losses: classification, region-text grounding, spatial emotion, and spatial confidence regression. At inference, LG-GER requires no detectors, no MLLM, and no multi-stream fusion, making GER practical for real-time and resource-constrained deployment. LG-GER has been evaluated on two benchmark GER datasets (GroupEmoW and GAF~3.0) and achieves competitive or superior results compared to state-of-the-art methods that require detection and multi-stream processing at inference.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑