发表机构
Wangxuan Institute of Computer Technology, Peking University; School of Computer Science, Wuhan University(北京大学王选计算机研究所; 武汉大学计算机学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ReART框架通过参考引导检索与AAS驱动的细化,在AffectiveArt 2026挑战赛Track 1中获第2,实现1.00的AAS与0.78总分,用于情感感知艺术生成。
AI 中文摘要
情感感知艺术图像生成要求模型同时满足语义内容、艺术风格和目标情感的需求。关键挑战在于艺术描述将这些维度合并为不明确的自由格式文本,使得笔触、构图和色调氛围等细粒度视觉属性难以具体锚定。我们提出ReART,一种参考引导的检索与细化框架。该方法将测试描述和EmoArt数据库中的每个图像标注分解为结构化视觉场,并在主体、布局、笔触和色调-情绪维度上进行逐场检索,以获取角色特定的视觉参考,提供仅文本描述无法传达的感知细节;这些参考与结构化提示一起用于初始合成。对于任何属性对齐分数(AAS)轴低于阈值的样本,AAS驱动的细化循环会诊断失败,构建指定保留元素、待修复错误和需避免操作的约束修复计划,按校正目的路由参考,并在结构保留约束下执行受控编辑。我们的系统在AffectiveArt 2026挑战赛的Track 1中排名第2,实现了1.00的完美AAS和0.78的总分。代码可在此https URL获取。
英文摘要
Emotion-aware artistic image generation requires a model to satisfy semantic content, artistic style, and target emotion simultaneously. The key challenge is that artistic captions conflate these axes into underspecified free-form text, making fine-grained visual attributes such as brushwork, composition, and tonal atmosphere difficult to ground concretely. We present ReART, a reference-guided retrieval and refinement framework. Our method decomposes test captions and each image annotation in the EmoArt database into structured visual fields, and performs field-wise retrieval over subject, layout, brush-line, and tone-mood dimensions to retrieve role-specific visual references that supply the perceptual detail text alone cannot convey; these references are used alongside a structured prompt for initial synthesis. For samples where any Attribute Alignment Score (AAS) axis falls below threshold, an AAS-driven refinement loop diagnoses failures, constructs constrained repair plans specifying elements to keep, errors to fix, and operations to avoid, routes references by correction purpose, and performs controlled editing under structural preservation constraints. Our system ranks 2nd in Track 1 of the AffectiveArt 2026 Grand Challenge, achieving a perfect AAS of 1.00 and an overall score of 0.78. Code is available at https://github.com/oceanflowlab/ReART.git.
CommentsAccepted by ACM Multimedia 2026 (Grand Challenge Track 1), 7 pages, 3 figures