arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EmoStyle:用于情感图像生成的风格专家情感调节

EmoStyle: Affective Conditioning of Style-Specialist Experts for Emotional Image Generation

Dexiang Hong, Yijie Guo, Weidong Chen, Xinyan Liu, Zixuan Zou, Zhendong Mao, Yongdong Zhang

arXiv 2607.10165首次发表:更新:

发表机构

University of Science and Technology of China; Harbin Institute of Technology(中国科学技术大学; 哈尔滨工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对情感感知艺术图像生成中训练与测试时属性不一致造成的控制差距问题,提出EmoStyle框架,利用语言模型推理器预测情感线索,编码情感字段指导生成,训练专用LoRA适配器结合风格表达情感,经视觉语言模型引导排序,在挑战赛中获佳绩。

AI 中文摘要

情感感知的艺术图像生成面临挑战,训练数据中的视觉和情感属性在测试时未明确提供,导致生成器在决定描绘内容及如何通过多种元素表达目标情感时存在控制差距。为此提出EmoStyle框架,先由语言模型推理器预测情感线索等,将情感字段编码为条件向量注入去噪块指导中间特征生成。因情感表达与风格有关,还为每种艺术风格训练专用LoRA适配器,最后通过轻量级视觉语言模型引导的候选选择步骤对生成图像排序。在2026年情感艺术挑战赛的Track 1中,相关团队提交成果获第一名。

英文摘要

Emotion-aware artistic image generation requires an image to match the input prompt, follow the specified artistic style, and convey the target emotion. In this challenge, the main difficulty is that the visual and affective attributes available in the training data are not explicitly provided at test time. Without these attributes, the generator has to decide not only what to depict, but also how the target emotion should be expressed through color, lighting, brushwork, composition, line, and layout. This creates a control gap between the available test prompt and the fine-grained conditions needed for emotion-aware artistic generation. To bridge this gap, we propose EmoStyle, a Z-Image-based framework that converts the input prompt into a structured generation state. An LLM reasoner first predicts affective cues (valence-arousal, dominant emotion, and therapeutic-effect labels) and an aspect-ratio decision. Instead of using these predictions only as additional prompt text, we encode the affective fields into an affective condition vector and inject it into the denoising blocks through AdaLN-style modulation. This allows the inferred control variables to directly guide the generation of intermediate features. Since emotional expression is also style-dependent, we further train a dedicated LoRA adapter for each artistic style bucket and select the corresponding expert during inference, enabling the same affective cues to be rendered with bucket-specific priors for color, texture, brushwork, and composition. Finally, a lightweight VLM-guided candidate selection step ranks the generated images based on prompt alignment, style consistency, emotional expression, and visual quality. In Track 1 of the AffectiveArt Challenge 2026, our USTC\_PI\_LAB\_TEAM submission achieved first place.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑