发表机构
School of Artificial Intelligence, Shenzhen Technology University; School of Innovation Design, Shenzhen Technology University; Stanford University(深圳技术大学人工智能学院; 深圳技术大学创新设计学院; 斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究情感图像编辑,EmoScope多智能体框架将任务转变为‘能编辑什么’,先发现可编辑空间,再平衡内容与情感,通过情感条件作用的可供性推理选择策略,在大规模评估中受青睐,还指出相关指标盲点及优势变化情况。
AI 中文摘要
情感图像编辑需要识别特定图像对目标情感的作用,而现有方法多基于预定义分类法等在有限策略空间内操作,常遗漏特定于图像和上下文的策略。我们引入EmoScope多智能体框架,将任务从‘如何编辑’转变为‘能编辑什么’。它先通过情感条件作用的可供性推理发现特定于图像的可编辑空间,再利用语义层次结构平衡内容一致性和情感表现力,最后执行并验证编辑。在大规模人类评估中,参与者平均有88.1%更青睐EmoScope。归因分析表明它选择适应目标情感的策略。最后,我们指出基于分类器的指标对非刻板、基于上下文的编辑存在情感条件盲点,并展示了EmoScope优势随图像 - 情感组合系统变化的相对内容 - 情感偏好 - 亲和力景观。
英文摘要
Emotional image editing requires more than applying affective filters or modifying predefined visual factors: an effective edit must identify what a particular image can afford for a target emotion. Existing affective image manipulation methods, including recent agentic variants, largely operate within bounded strategy spaces based on predefined factor taxonomies, knowledge libraries, or conventional editing templates, and therefore often miss image-specific, context-grounded strategies. We introduce EmoScope, a multi-agent framework that reframes the task from "how should I edit?" to "what can I edit?" EmoScope first discovers an image-specific editable space through emotion-conditioned affordance reasoning, then uses a semantic hierarchy of anchors, variables, and context to balance content consistency and emotional expressiveness before executing and verifying the edit. Because its plans are expressed as image-specific affordances rather than retrieved templates, EmoScope also exposes the editing strategy as an interactive surface for user refinement at the plan level. In a large-scale human evaluation covering all eight Mikels emotion categories, with 4,693 valid responses across 1,824 pairwise questions, participants preferred EmoScope over two competitive baselines by 88.1% on average. Attribution analysis further shows that EmoScope selects target-emotion-adaptive strategies rather than applying a uniform template. The same affordance-level plan also supports lightweight user refinement in an interactive pilot. Finally, we show that classifier-based metrics exhibit emotion-conditional blind spots toward non-stereotypical, context-grounded edits, and present a relative content-emotion preference-affinity landscape showing that EmoScope's advantage varies systematically across image-emotion combinations.