arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EmoWorld:用于可控情感视频生成的解耦情感场

EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation

Bingyuan Wang, Baistan Zhyldyzbekov, Kunyu Feng, Zeyu Wang

arXiv 2608.06231首次发表:更新:

AI 中文总结

EmoWorld是一种解耦情感场框架,通过VAS、SAS、TAS三个引导模块优化可控情感视频生成,在Wan2.2等基准上实现情感对齐、线索检测及过渡单调性提升,具备多骨干网络可移植性。

AI 中文摘要

情感影响观众对场景的理解,但现有视频生成器会将全局氛围、带情感的语义线索与时间进程纠缠在单一文本条件中。本文提出EmoWorld,这是一种在冻结的流匹配视频扩散Transformer(Video DiT)内解耦这些因素的框架。一次性准备阶段从保持几何的中性全景图和经情感编辑的全景图中提取特定层的情感方向和可复用线索库。推理时,视觉氛围引导(VAS)将氛围方向注入隐藏状态,语义情感引导(SAS)分离出可单独扩展的提示残差以处理语义线索,时间情感引导(TAS)在去噪和视频时间内插值端点残差场。在Wan2.2上,VAS使目标情感对齐度提升19%,同时降低时间波动代理值48%;SAS使目标情感对齐度提升37%,检测到的带情感线索数量增加36%;TAS使过渡单调性提升15%,优于最强基线。EmoWorld在文本到视频和图像到视频设置中针对27种情感类别进行评估,展现出在多个Video DiT骨干网络间的可移植性,且支持相机条件下的合成,无需更新生成器参数。

英文摘要

Emotion shapes how viewers interpret a scene, yet existing video generators entangle global atmosphere, affect-bearing semantic cues, and temporal progression within a single text condition. We present EmoWorld, a framework that decouples these factors within a frozen flow-matching video diffusion transformer (Video DiT). A one-time preparation stage extracts layer-specific affect directions and a reusable cue library from geometry-preserving neutral and emotion-edited panoramas. At inference, Visual Atmosphere Steering (VAS) injects atmosphere directions into hidden states, Semantic Affective Steering (SAS) isolates a separately scalable prompt residual for semantic cues, and Temporal Affective Steering (TAS) interpolates endpoint residual fields across denoising and video time. On Wan2.2, VAS improves target-emotion alignment by 19% while reducing a temporal-fluctuation proxy by 48%; SAS improves target-emotion alignment by 37% and increases detected affect-bearing cues by 36%; and TAS improves transition monotonicity by 15% over the strongest baseline. EmoWorld is evaluated across 27 emotion categories in text-to-video and image-to-video settings, demonstrates portability across multiple Video-DiT backbones, and supports camera-conditioned composition without updating generator parameters.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑