发表机构
Shanghai Jiao Tong University; CreativeFitting(上海交通大学; CreativeFitting)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大规模短剧生成的视觉连续性瓶颈,提出无训练的SEAM记忆图,在SEAM-Bench上提升连续性召回率,部署后获96.5%导演接受率。
AI 中文摘要
短剧生成已发展为大型工业化流水线,当它从孤立镜头扩展至剧集层面时,视觉连续性成为关键瓶颈。现有智能体框架独立生成每个镜头,导致镜头间的上下文漂移,道具、角色姿势与调度变得不一致,这些细微偏差在拼接后会放大为严重的视觉断裂。我们提出SEAM(Shot Entity-Attribute Memory,镜头实体-属性记忆),这是一种无训练、与模型无关的记忆图,完全在提示文本层修复连续性:为每个镜头提取多维状态,在生成的图上仅检索因果先验上下文,选择性过滤后通过自然语言提示重写注入保留的约束。我们还发布了SEAM-Bench,这是一个双盲连续性故事板基准,在该基准上,SEAM将跨剧集连续性召回率从0.700提升至0.946,可在6种主流文本模型间泛化,在生成图像层产生一致但尚未显著的提升。作为CreativeFitting的SEAM-Agent生产流水线的强制阶段部署于201个镜头时,SEAM达到96.5%的导演接受率,且无不安全注入;保守反事实分析显示,该接受率至少有21.9个百分点归因于其跨剧集记忆。
英文摘要
Short-drama generation has grown into a large, industrialized pipeline, and as it scales from isolated shots to the episode level, visual continuity has become a critical bottleneck. Current agent frameworks generate each shot in isolation, so context drifts across shots and props, character posture, and blocking turn inconsistent. Once assembled, these small discrepancies amplify into severe visual breaks. We present SEAM (Shot Entity-Attribute Memory), a training-free, model-agnostic memory graph that repairs continuity entirely at the prompt-text layer by extracting a multi-dimensional state for every shot, retrieving only causally prior context over the resulting graph, filtering it selectively, and injecting the surviving constraints by natural-language prompt rewriting. We further release SEAM-Bench, a double-blind continuity storyboarding benchmark, on which SEAM raises cross-episode continuity recall from 0.700 to 0.946, generalizes across six mainstream text models, and yields consistent, though not yet significant, gains at the generated-image layer. Deployed as a mandatory stage in CreativeFitting's SEAM-Agent production pipeline over 201 shots, SEAM reaches a 96.5% director-acceptance rate with zero unsafe injections; a conservative counterfactual attributes at least 21.9 percentage points of that rate to its cross-episode memory.
CommentsPreprint. 12 pages, 6 figures