arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12830cs.CV

通过基于效价-唤醒锚定的强化学习在图像生成中平衡情感对齐与语义一致性

Balancing Emotional Alignment and Semantic Consistency in Image Generation via Reinforcement Learning with Valence-Arousal Anchoring

Jisheng Dang, Zhenxuan Wang, Bin Li, Ronghao Lin, Bin Hu, Tat-Seng Chua

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出锚定正则化的Flow-GRPO框架,结合连续效价-唤醒条件与GRPO及中性语义锚,在3,300个组合上降低情感误差并提升CLIPScore,平衡情感对齐与语义一致性。

中文摘要 AI 辅助

文本到图像生成中的连续情感控制要求模型在不改变提示词所描述的对象、布局或场景的前提下提升情感对齐度。现有的监督式情感注入方法通常优化特征空间代理,因此可能表现出情感-语义漂移,即更强的情感条件作用伴随非预期的内容变化。我们通过一个流匹配图像生成框架来解决这一问题,该框架结合了连续效价-唤醒(VA)条件、组相对策略优化(GRPO)以及一个中性语义锚。确定性概率流ODE被转换为保持边缘分布的SDE,从而为轨迹采样和策略比率估计提供非退化转移密度。一个冻结的基于CLIP的VA回归器提供终端奖励,衡量预测VA坐标与目标VA坐标之间的距离,而由同一提示词在零VA条件下生成的图像则为语义保持提供特征空间参考。在线RL采样使用简化的去噪调度,而推理时保留原始调度。在3,300个提示-情感组合上的实验显示,与VA条件基线相比,效价和唤醒误差显著降低,并且相对于EmotiCrafter,CLIPScore有所提高,但在无参考图像质量方面存在可衡量的权衡。结果支持锚定正则化的Flow-GRPO作为在连续情感图像合成中平衡情感对齐与语义一致性的实用方法。

英文摘要

Continuous emotion control in text-to-image generation requires a model to improve affective alignment without changing the objects, layout, or scene described by the prompt. Existing supervised emotion-injection methods often optimize feature-space proxies and may therefore exhibit emotion-semantic drift, in which stronger emotional conditioning is accompanied by unintended content changes. We address this problem with a flow-matching image-generation framework that combines continuous valence-arousal (VA) conditioning, Group Relative Policy Optimization (GRPO), and a neutral semantic anchor. The deterministic probability-flow ODE is converted into a marginal-preserving SDE, yielding non-degenerate transition densities for trajectory sampling and policy-ratio estimation. A frozen CLIP-based VA regressor supplies a terminal reward measuring the distance between the predicted and target VA coordinates, while an image generated from the same prompt under zero VA conditioning provides a feature-space reference for semantic preservation. A reduced denoising schedule is used for online RL sampling, whereas the original schedule is retained at inference. Experiments on 3,300 prompt-emotion combinations show substantially lower valence and arousal errors than the VA-conditioned baseline and an improved CLIPScore relative to EmotiCrafter, with a measurable trade-off in reference-free image quality. The results support anchor-regularized Flow-GRPO as a practical approach to balancing emotional alignment and semantic consistency in continuous-affect image synthesis.

发表机构

  • Lanzhou University(兰州大学)
  • China University of Mining and Technology(中国矿业大学)
  • Shenzhen University(深圳大学)
  • Beijing Institute of Technology(北京理工大学)
  • National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑