发表机构
Nanyang Technological University; Institute for Infocomm Research, A*STAR; Ant Group; Nanjing University of Information Science and Technology(南洋理工大学; 资讯通信研究院,新加坡科技研究局; 蚂蚁集团; 南京信息工程大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DiVA提出一种结合多模态大语言模型与堆叠视频流水线的数字生命模拟框架,通过锚定视频延续模块确保长期交互的视觉质量与连贯性,显著优于现有方法。
AI 中文摘要
我们提出了DiVA,一个深度交互的数字生活模拟器,开创了数字角色世界中长期、开放式交互体验的新范式。DiVA的架构将多模态大语言模型(MLLM)作为路由器,与精心设计的堆叠视频流水线相结合,实现带有动作和音频响应的无缝多轮交互。为了保持连续性并避免性能退化,我们将生成建模为一个三部分耦合系统:等待视频、动作视频以及它们之间的过渡。这些过渡由我们的锚定视频延续(AVC)模块关键处理,该模块将角色恢复到稳定状态以防止退化。通过编码来自前一个动作视频片段的信息,AVC确保了平滑过渡,显著减少了当前视频过渡方法中常见的相机抖动和不一致性。这种设计还实现了音频驱动模型通常难以处理的复杂姿态变化(例如从坐姿到站姿)。这些系统设计共同确保了长时间体验中的高保真身份、连贯性和动态性。为了验证我们的流水线设计,我们通过用主流的长视频、延续和插值方法替换我们的核心生成模块,将我们的系统与替代方案进行了全面比较。我们进一步分析了三阶段设计、锚点状态选择、过渡自然性、空间接地以及质量-延迟权衡的必要性,并将比较扩展到额外的长格式音频驱动虚拟形象模型。结果证实,DiVA在维持长期视觉质量和逼真度方面显著优于其他方法,验证了其作为可持续交互模拟的有效性。
英文摘要
We present DiVA, a deeply interactive digital life simulator pioneering a new paradigm for long-term, open-ended interactive experiences within digital character worlds. DiVA's architecture pairs a Multimodal Large Language Model (MLLM) as a router with a meticulously designed stacked video pipeline for seamless, multi-turn interactions with action and audio response. To maintain continuity and avoid degradation, we model generation as a three-part coupled system: waiting video, action video, and the transitions between them. These transitions are critically handled by our Anchored Video Continuation (AVC) module, which returns the character to stable states to prevent degradation. By encoding information from the preceding action video segment, AVC ensures smooth transitions, significantly reducing camera jitter and inconsistencies common in current video transition methods. This design also enables complex pose changes (e.g., sitting to standing) typically difficult for audio-driven models. These system designs together ensure high-fidelity identity, coherence, and dynamics for extended experiences. To validate our pipeline design, we comprehensively compare our system against alternatives by replacing our core generation module with mainstream long-video, continuation, and interpolation methods. We further analyze the necessity of the three-stage design, anchor-state selection, transition naturalness, spatial grounding, and the quality-latency trade-off, and we expand the comparison to additional long-form audio-driven avatar models. Results confirm DiVA is markedly superior in maintaining long-term visual quality and realism, validating its effectiveness as a sustainable, interactive simulation.
CommentsProject page: https://steller-cheng.github.io/DiVA/