Omni-Perception Policy Optimization for Multimodal Emotion Reasoning
全模态感知策略优化用于多模态情感推理
机构 * University of Science and Technology of China(中国科学技术大学) ; SenseTime Research(感时间研) ; National University of Singapore(新加坡国立大学) ; Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)
AI总结 提出OPPO强化学习框架,通过全模态感知奖励和损失优化多模态感知,提升情感推理中模态利用率和忠实度,在多个基准上达到最优。
Comments Accepted at ICML 2026. Project page: https://zjycutieee.github.io/OPPO-page/ Code: https://github.com/ZhiyuanHan-Aaron/OPPO