发表机构
Nanjing University; Tencent YoutuLab; The Hong Kong University of Science and Technology(南京大学; 腾讯优图实验室; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对艺术情感理解中多子任务权衡问题,提出多智能体框架ArtSociety,通过异构专家协作与免训练控制器,在官方测试集上取得0.8870总体得分,并揭示协作与数据质量是关键。
AI 中文摘要
AffectiveArt多维艺术情感理解任务要求联合预测艺术品细粒度情感(12类,头尾比1549:1)、二元效价/唤醒度以及五个属性基础描述——这些子任务表现出强烈的经验性权衡,因此我们尝试的单一模型解决方案无法同时优化所有子任务。我们提出ArtSociety,一个多智能体框架,它组装了异构多模态专家——一个DINOv2-Giant视觉智能体(A1)、一个场景基础CoT微调MLLM(A2)和三个闭源推理智能体(A3-A5)——并通过两个免训练控制器协调它们:(i)一个稀有类别感知投票仲裁器,降低尾部情感的一致性阈值,利用跨智能体家族的去相关错误模式;以及(ii)一个描述优先推理智能体,其DESCRIBE-then-CLASSIFY思维链在标签承诺前强制视觉证据,产生近乎完美的基于属性的描述。任务路由策略将困难的情感任务分配给完整的五智能体集成,而将接近饱和的效价/唤醒度和生成描述任务分配给单个最强的推理智能体。在官方测试集(1000件艺术品)上,ArtSociety实现了0.8870的总体得分(分类0.7789,描述0.9952)。一项包含十一个变体的消融研究表明,一旦方法和规模在0.76左右饱和,决定性收益来自智能体协作和数据侧监督——在旧数据上训练的30B MoE模型并不优于在更好数据上训练的8B模型。代码可在https此URL获取。
英文摘要
The AffectiveArt Multidimensional Art Emotion Understanding task asks to jointly predict an artwork's fine-grained emotion (12 classes, 1549:1 head-to-tail ratio), binary valence/arousal, and five attribute-grounded descriptions -- sub-tasks that exhibit strong empirical trade-offs, so the single-model solutions we tried do not jointly optimize all of them well. We present ArtSociety, a multi-agent framework that assembles heterogeneous multimodal experts -- a DINOv2-Giant vision agent (A1), a scene-grounded CoT fine-tuned MLLM (A2), and three closed-source reasoning agents (A3-A5) -- and coordinates them with two training-free controllers: (i) a rare-class-aware voting arbiter that lowers the agreement threshold for tail emotions, exploiting decorrelated error patterns across agent families; and (ii) a description-first reasoning agent whose DESCRIBE-then-CLASSIFY chain of thought forces visual evidence before label commitment, yielding near-perfect grounded descriptions. A task-routing policy directs the hard emotion task to the full five-agent ensemble while assigning the near-saturated valence/arousal and generative description tasks to the single strongest reasoning agent. On the official test set (1,000 artworks), ArtSociety achieves an Overall Score of 0.8870 (Classification 0.7789, Description 0.9952). An eleven-variant ablation study reveals that, once method and scale saturate at around 0.76, the decisive gains come from agent collaboration and data-side supervision -- a 30B MoE model trained on older data does not outperform an 8B model trained on better data. Code is available at https://github.com/swordlidev/ArtSociety
CommentsAccepted at the ACM Multimedia 2026 Grand Challenge (AffectiveArt). 8 pages, 5 figures, 3 tables. Code: https://github.com/swordlidev/ArtSociety