arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21013cs.AIcs.CV

EmoAgent-R1:基于强化学习的动态智能体专业化实现多模态情感理解

EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization

发表机构广东工业大学 · 南方科技大学 · 西安交通大学
另 1 家 · 查看机构详情
  • Guangdong University of Technology(广东工业大学)
  • Southern University of Science and Technology(南方科技大学)
  • Xi’an Jiaotong University(西安交通大学)
  • Guizhou Provincial Laboratory of Big Data, State Key Laboratory of Public Big Data, Guizhou University(贵州大学大数据省级重点实验室、公共大数据国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

Lihuang Fang, Yuchen Zou, kebing Jin, Jinghui Qin

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对现有多模态情感识别方法忽视情感源动态性和复杂性的问题,提出基于强化学习的动态智能体专业化框架EmoAgent-R1,采用冷启动策略和两步智能体工作流程训练,并用P-GRPO优化,实验证明其在情感推理性能和优化稳定性上更优。

中文摘要 AI 辅助

多模态大语言模型在多模态情感识别任务中取得了显著性能,将其提升到了复杂情感理解的新高度。然而,现有基于多模态大语言模型的方法常使用固定提示来感知情感,忽略了多模态输入中情感源的动态性和复杂性。为解决这些问题,我们提出了基于强化学习的动态智能体专业化框架(EmoAgent-R1),通过基于强化学习的动态智能体专业化来优化多模态大语言模型的情感识别、推理和泛化能力。具体而言,首先采用冷启动策略,通过用合成答案条件思维链数据和智能体路由数据训练,赋予多模态大语言模型初步情感识别、推理和智能体路由能力。然后,通过强化学习在两步智能体工作流程中训练多模态大语言模型以感知情感,包括智能体选择和智能体专业化。为有效训练EmoAgent-R1,我们提出了新颖的渐进式组相对策略优化(P-GRPO),将基于组的相对优势与受PMI启发的渐进式令牌级调制相结合,将稀疏奖励转化为细粒度学习信号,减轻GRPO中的粗粒度均匀信用分配问题。在多模态情感识别基准上的大量实验证明了EmoAgent-R1在更强的情感推理性能和改进的优化稳定性方面的优越性。

英文摘要

Multimodal large language models (MLLMs) have achieved impressive performance in multimodal emotion recognition (MER) tasks and lifted MER to a new level that is complex emotion understanding with advanced video understanding abilities and natural language description. However, existing MLLM-based methods often use a fixed prompt to perceive the emotions, ignoring the dynamicity and complexity of the emotion source in the multimodal inputs. To address these issues, we propose a novel Reinforcement Learning-based Dynamic Agent Specialization framework (\textbf{EmoAgent-R1}) to optimize the emotion recognition, reasoning, and generalization abilities of an MLLM with dynamic agent specialization based on reinforcement learning. Specifically, we first adopt a cold start strategy to endow an MLLM with preliminary emotion recognition, reasoning, and agent routing ability by training with synthetic answer-conditioned chain-of-thought data and agent routing data. Then, we further train the MLLM with reinforcement learning to perceive emotions in a two-step agentic workflow with agent selection and agent specialization. To effectively train EmoAgent-R1, we propose a novel Progressive Group-Relative Policy Optimization (P-GRPO) to combine group-based relative advantages with a PMI-inspired progressive token-level modulation to transform sparse rewards into fine-grained learning signals, mitigating the coarse-grained uniform credit assignment issue in GRPO. Extensive experiments on MER benchmarks demonstrate the superiority of our EmoAgent-R1 in stronger emotion reasoning performance and improved optimization stability.

↑