arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AffectOmni:面向社交与艺术相关场景的可通过强化学习验证的以人为中心的接地情感推理

AffectOmni: RL-Verifiable People-Centric Grounded Affective Reasoning for Social and Art-Related Scenes

Yibo Wang, Rui Yang, Jisheng Dang, Bimei Wang, Yitao Wu, Pengfei Cao, Wencan Zhang, Hong Peng, Bin Hu, Tat-Seng Chua

arXiv 2608.26193首次发表:更新:

发表机构

Lanzhou University; Hainan University; National University of Singapore(兰州大学; 海南大学; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AffectOmni是面向社交与艺术场景的可验证情感推理框架,通过GRPO训练,引入两类奖励优化以人为中心的证据选择与时间推理,经多数据集实验较7B基线实现显著性能提升。

AI 中文摘要

多模态大语言模型(MLLM)在视觉问答(VQA)和场景理解任务中表现出色,但情感推理仍易受捷径行为影响。模型可能预测出正确答案,却忽略微表情、肢体语言等以人为中心的线索,这削弱了可追溯性和外部验证性。现有强化学习方法主要对上下文或逻辑一致性进行奖励,未明确要求关注人类证据;此外,“LLM作为评判者”的评分常出现分数聚类问题,降低了奖励的区分度。我们提出AffectOmni,这是一个经GRPO训练的可验证情感推理框架。AffectOmni引入以人为中心证据选择和时间结构化推理的“以人为中心”与“时间顺序”奖励,采用组内对比评分以生成更稳定、区分度更高的奖励信号。为实现验证,“思考摘要器”将自由形式的理由转换为可执行的证据指令,通过SAM3接地到像素级证据区域,为训练循环外提供可外部审计的接口。在IntentBench、Daily Omni和WorldSense上的实验显示,与开源7B规模基线相比,AffectOmni取得了一致的性能提升,其中情感识别任务提升4.66%,时间敏感任务提升14.29%。代码可在该https地址获取。

英文摘要

Multimodal large language models (MLLMs) achieve strong performance on VQA and scene understanding, yet affective reasoning remains vulnerable to shortcut behavior. Models may predict correct answers while neglecting people-centric cues such as micro expressions and body language, which weakens traceability and external verification. Prior reinforcement learning approaches mainly reward context or logical coherence without explicitly enforcing attention to human evidence. In addition, LLM as a Judge scoring often suffers from score clustering, which reduces reward discriminability. We propose AffectOmni, a GRPO trained framework for verifiable affective reasoning. AffectOmni introduces People Focus and Temporal Order rewards to encourage people-centric evidence selection and temporally structured reasoning, and it adopts within-group comparative scoring to produce more stable and discriminative reward signals. For verification, a Thinking Summarizer converts free form rationales into executable evidence instructions, which are grounded into pixel level evidence regions via SAM3 to provide an externally auditable interface outside the training loop. Experiments on IntentBench, Daily Omni, and WorldSense show consistent improvements over open source 7B scale baselines, including gains of 4.66% on emotion recognition and +14.29% on temporally sensitive tasks. Code is available at https://github.com/eliot127825-rgb/AffectOmni_nobody.

Comments12 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑