AffectDelta:超越情感标签的图像编辑
AffectDelta: Beyond Emotion Labels for Image Editing
浏览论文内容
中文总结 AI 辅助
AffectDelta是一种将图像编辑视为八维情感分布过渡的感知源图像编辑器,通过构建AffectPair-249K数据集训练,在情感对齐和内容保留上优于基线方法。
中文摘要 AI 辅助
情感驱动的图像编辑旨在通过修改源图像中与情感相关的视觉线索,唤起指定的目标情感,同时保留原始场景的整体构图和语义结构一致性。现有的场景级编辑器通常用单一情感类别指定目标,且往往从操作级文本指令中学习视觉变换。一个类别会将混合情感终点归为一个主导标签,而语言无法精确量化共存情感应如何增加、减少或保持稳定。我们提出AffectDelta,这是一种感知源图像的编辑器,将编辑视为八维情感分布之间的过渡。一个冻结的情感分布预测器(Emotion Distribution Predictor)估计源状态,而带符号的源到目标的差值则编码了请求过渡的方向和幅度。在AffectDelta内部,一个内部过渡编码器和感知源图像的扩散骨干网络(source-aware diffusion backbone)共同将该信号转换为依赖上下文的语义和外观变化。为训练该模型,我们构建了AffectPair-249K数据集,包含248,841个源-目标对,这些对带有预测的八维分布,涵盖跨类别和类别内过渡。针对六个基线的实验,结合定量评估与定性比较,表明该方法在情感对齐和内容保留方面有所改进,而消融实验验证了我们的设计选择。代码和数据集将在接收后公开提供。
英文摘要
Emotion-driven image editing aims to evoke a specified target emotion by modifying emotion-relevant visual cues in a source image, while preserving the overall composition and semantic-structural coherence of the original scene. Existing scene-level editors typically specify the target with a single emotion category and often learn visual transformations from operation-level text instructions. A category collapses a mixed affective endpoint into one dominant label, while language cannot precisely quantify how coexisting emotions should increase, decrease, or remain stable. We introduce AffectDelta, a source-aware editor that treats editing as a transition between eight-dimensional emotion distributions. A frozen Emotion Distribution Predictor estimates the source state, and the signed source-to-target difference encodes the direction and magnitude of the requested transition. Within AffectDelta, an internal transition encoder and a source-aware diffusion backbone jointly translate this signal into context-dependent semantic and appearance changes. To train this formulation, we construct AffectPair-249K, comprising 248,841 source-target pairs with predicted eight-dimensional distributions and spanning both cross-category and within-category transitions. Experiments against six baselines, combining quantitative evaluation with qualitative comparisons, demonstrate improved affective alignment and content preservation, while ablations validate our design choices. Code and dataset will be made publicly available upon acceptance.
发表机构
- Vanderbilt University(范德堡大学)
- Research Institute of Electrical Communication, Tohoku University(东北大学电气通信研究所)
机构由 AI 辅助整理,请以论文原文为准。