我们真的需要参数超过10亿的多模态情感语言模型吗?
Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters?
浏览论文内容
中文总结 AI 辅助
质疑多模态情感识别需参数超10亿模型的假设,提出轻量级MER框架Light-MER,通过知识蒸馏及两种优化策略,在九个基准数据集实验中实现最优性能并显著提高推理效率,凸显小模型潜力。
中文摘要 AI 辅助
多模态大语言模型的进展显著提升了多模态情感识别(MER)性能,并能通过联合建模视频、音频和语言等实现可解释描述生成。然而,性能提升常伴随模型参数规模增大(如至少70亿),带来高计算成本和低推理效率,阻碍在资源受限平台实时部署。本文质疑大模型必要性,提出轻量级MER框架Light-MER,通过知识蒸馏实现更好更快的多模态情感理解与识别。具体介绍两种优化策略:结合切片瓦瑟斯坦距离与隐藏状态对齐的最优传输损失,以及基于GRPO平衡MER性能和效率的多奖励优化策略。在九个基准数据集上的大量实验表明,Light-MER实现了最优性能且显著提高推理效率,凸显小多模态情感语言模型的潜力。
英文摘要
Recent advances in multimodal large language models (MLLMs) have significantly improved the performance of multimodal emotion recognition (MER) and enabled interpretable description generation by jointly modeling video, audio, and language, etc. However, these performance improvements are often accompanied by an increase in model parameter size (e.g, at least 7B), which simultaneously incurs high computational costs and reduces inference efficiency, thereby hindering real-time deployment on resource-constrained platforms such as robots and mobile devices. This raises a fundamental question: do we really need the multimodal MER model larger than 1B parameters for high-quality MER? In this paper, we challenge the assumption that larger models are inherently necessary and proposes a lightweight MER framework (called Light-MER), which achieves better and faster multimodal sentiment understanding and recognition through knowledge distillation. It can transfer knowledge from a strong, large-scale teacher model to a lightweight sub-billion-parameter student model, aiming to preserve rich multimodal emotion reasoning and recognition while substantially improving deployment efficiency. Specifically, we introduce two new optimization strategies to enhance knowledge transfer: (1) a new optimal transport loss that combines Sliced Wasserstein Distance with hidden-state alignment, and (2) a new multi-reward optimization strategy based on GRPO that balances MER performance and efficiency, aimed at further enhancing the learning capabilities of student models. Extensive experiments on nine benchmark datasets demonstrate that Light-MER achieves state-of-the-art performance while significantly improving inference efficiency. This highlights the strong potential of small multimodal emotion language models for future research. Code is available at https://github.com/GAIR-Lab/Light-MER.
发表机构
- University of Glasgow(格拉斯哥大学)
- Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
- School of Artificial Intelligence, Shandong University(山东大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。