arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Sidecar:用于保持角色一致性的自由形式视觉叙事的无需训练的语义复用

Sidecar: Training-Free Semantic Reuse for Character-Consistent Free-form Visual Storytelling

Sibo Dong, Sarah Adel Bargal

arXiv 2608.27280首次发表:更新:

发表机构

Georgetown University(乔治城大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对自由形式视觉叙事中角色一致性难维持的问题,提出无需训练的即插即用语义增强模块Sidecar,可提升提示-图像对齐度与角色一致性,且计算开销小。

AI 中文摘要

视觉叙事需要生成遵循叙事逻辑且在各帧中保持角色身份一致的图像。在自由形式故事生成中,角色仅在首次引入时被完整描述,后续仅通过类型提及或代词指代,虽更贴近自然叙事,但后续提示可能遗漏重要的身份相关语义,使角色一致性更难维持。我们提出Sidecar,即一个即插即用的语义增强模块,它保留初始描述中的实体级信息,并将缺失的语义注入后续的提示嵌入中。Sidecar无需额外训练,也不修改基础扩散模型的架构。在FreeStoryBench上的实验表明,Sidecar在多个基于SDXL和FLUX的基线模型上,持续提升了提示-图像对齐度和角色一致性,且计算开销可忽略不计。

英文摘要

Visual storytelling requires generating images that follow a narrative while preserving consistent character identities across frames. In free-form story generation, a character is fully described only when first introduced and is later referred to by a type-level mention or pronoun. Although this setting better reflects natural storytelling, later prompts may omit important identity-related semantics, making character consistency more difficult to maintain. We propose \textbf{Sidecar}, a plug-and-play semantic augmentation module that preserves entity-level information from the initial description and injects the missing semantics into later prompt embeddings. Sidecar requires no additional training and does not modify the architecture of the base diffusion model. Experiments on FreeStoryBench show that Sidecar consistently improves prompt-image alignment and character consistency across multiple SDXL- and FLUX-based baselines, with negligible computational overhead.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑