arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TaoMate:用于实时音频-视频数字人生成的锚点引导内存桥接演化和参考状态

TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation

Qijun Gan, Chenwei Zhang, Meiguang Jin, Junfeng Ma, Qiu Shen

arXiv 2607.24359首次发表:更新:

AI 中文总结

研究实时长格式数字人生成问题,提出TaoMate框架,通过锚点引导、内存桥接等方法,在保持稳定外观和视听同步下,实现少步联合音频-视频生成,且能跨块并行执行加速自回归推理。

AI 中文摘要

实时长格式数字人生成依赖因果模型来扩展视听内容,同时保持主体外观和跨连续片段的视听同步。有界缓存保留局部运动和语音上下文,但会丢弃旧证据,而关注完整的生成历史计算成本高且会传播累积错误。我们提出了TaoMate,一种用于少步联合音频-视频生成的锚点引导持久内存框架。该框架保留不可变视觉锚点,将完整的视频和音频块压缩为固定容量的动态状态,并通过特定模态的残余注意力检索这些状态,而无需扩展活动缓存。一种参考感知调制方法还根据动态和锚点外观统计对视频特征进行条件设置。锚点保留因果上下文蒸馏在保持不可变视觉锚点不受干扰的同时,改变展开范围、前缀来源和缓存历史可靠性。通过将持久内存与阶段局部去噪依赖性分离,TaoMate进一步允许跨块进行阶段并行执行,加速自回归推理而无需特定于管道的重新训练。我们用外观、时间、同步、面部和语音诊断评估长格式视频延续。结果表明,TaoMate在自回归生成下在提示条件片段中保持稳定外观和强大的视听同步。我们的项目页面是这个https URL。

英文摘要

Real-time long-form digital-human generation relies on causal models to extend audio-visual content while preserving subject appearance and audio-video synchronization across successive segments. A bounded cache retains local motion and phonetic context but discards older evidence, whereas attending to the complete generated history is computationally expensive and can propagate accumulated errors. We present \method, an anchor-guided persistent-memory framework for few-step joint audio-video generation. The framework preserves an immutable visual anchor, compresses completed video and audio blocks into fixed-capacity dynamic states, and retrieves those states through modality-specific residual attention without extending the active cache. A reference-aware modulation method additionally conditions video features on dynamic and anchor appearance statistics. Anchor-preserving causal-context distillation varies rollout horizon, prefix provenance, and cache-history reliability while keeping the immutable visual anchor unperturbed. By separating persistent memory from stage-local denoising dependencies, \method further admits stage-parallel execution across blocks, accelerating autoregressive inference without pipeline-specific retraining. We evaluate long-form video continuations with appearance, temporal, synchronization, facial, and speech diagnostics. Results show that \method preserves stable appearance across prompt-conditioned segments and strong audio-visual synchronization under autoregressive generation. Our project page is https://taoliveaigc.github.io/TaoMate.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑