面向持续故事与交互世界的长时序视听生成
Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
浏览论文内容
中文总结 AI 辅助
研究针对长时序视听生成的需求,提出JoyAI-Echo-1.5系统,含长视频与世界模型变体,通过专用技术实现跨镜头一致性等性能,在相关基准上取得领先结果,为生成连贯内容提供基础。
中文摘要 AI 辅助
视频生成正从孤立片段向长形式叙事与交互世界发展,要求模型能保留实体身份、遵循用户控制并在长期生成中保持稳定。我们提出JoyAI-Echo-1.5,这是一个统一的视听生成系统,包含两个专用变体:长视频变体引入可组合的跨镜头记忆,该记忆聚合多个先前镜头的视觉证据及经全镜头音频过滤得到的说话人线索,能在文本、图像和记忆条件的灵活组合下实现持续的角色外观与语音身份;世界模型变体将异构导航输入转换为校准的度量6自由度(6-DoF)相机轨迹,并通过几何感知的条件通路注入,支持跨灵活视点的控制器无关交互。为支持高效的长时序生成,我们采用渐进式教师强制及在自生成回放上的短、长时序自梯度强制,将双向视听骨干网络转换为因果少步生成器。实验在两种设置下均表现出优异性能:JoyAI-Echo-1.5在跨镜头一致性、视觉质量、文本对齐和语音保真度上优于现有长视频基线;其世界模型变体在WBench上排名第一,平均得分81.7,并在SANA-WM-Bench上实现领先的视觉质量与长时序持续性。这些结果共同表明,记忆、几何控制和感知回放的训练为生成连贯故事与持续演化的交互世界提供了实用基础。项目页面:this https URL
英文摘要
Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain stable over extended rollouts. We present JoyAI-Echo-1.5, a unified audio-visual generation system with two purpose-built variants. The long-video variant introduces composable cross-shot memory that aggregates visual evidence across multiple prior shots and speaker cues derived from speech-filtered full-shot audio, enabling persistent character appearance and voice identity across flexible combinations of text, image, and memory conditioning. The world-model variant converts heterogeneous navigation inputs into calibrated metric 6-DoF camera trajectories and injects them through a geometry-aware conditioning pathway, enabling controller-agnostic interaction across flexible viewpoints. To support efficient long-horizon generation, we transform a bidirectional audio-visual backbone into a causal few-step generator using progressive teacher forcing and short- and long-horizon Self-Gradient Forcing on self-generated rollouts. Experiments demonstrate strong performance in both settings. JoyAI-Echo-1.5 achieves improvements over existing long-video baselines in cross-shot consistency, visual quality, text alignment, and speech fidelity. Its world-model variant ranks first on WBench, with an average score of 81.7, and achieves leading visual quality and long-horizon persistence on SANA-WM-Bench. Together, these results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds. Project page: https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/.
发表机构
- Joy Future Academy, JD(京东探索研究院)
机构由 AI 辅助整理,请以论文原文为准。