arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11752cs.CVcs.SD

UniSwap:用于说话视频的流音频-视觉身份交换

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos

Yuxuan Zhang, Haozhong Xiong, Jiayi Song, Jinpeng Yu, Yang Shi, Jiaming Liu, Ruihua Huang, Liwei Wang

首次发表
浏览论文内容

中文总结 AI 辅助

UniSwap是首个流音频-视觉身份替换框架,通过多阶段适配技术解决现有方法音视频一致性差等问题,实现高效稳定的说话视频身份替换。

中文摘要 AI 辅助

说话视频的角色替换需要在保留源视频的运动、场景、语言内容及音视频时序的同时,协调传递外观与语音。现有方法采用分别优化的模型处理两种模态,难以保证音视频一致性。本文提出UniSwap,首个用于说话视频的流音频-视觉身份替换框架。给定源视频、参考图像及参考语音片段,UniSwap在单个音频-视觉扩散Transformer中传递参考外观与音色,同时保留源内容与动态。针对对齐跨身份训练对稀缺的问题,本文引入交换-重建流程:从真实片段中移除视觉与语音身份,并将原始片段作为重建目标。基于双向骨干,本文通过三阶段逐步适配模型:上下文预训练用于联合替换、条件流适配用于块因果KV缓存生成、高效自强迫DMD用于缓解暴露偏差并将每块采样去噪步骤从30步降至3步;高效多LoRA切换使三个DMD角色共享单个冻结骨干;特征旋转位置编码分解将缓存位置保持在训练范围内,支持稳定的长视频推理。实验表明,UniSwap具备出色的音视频同步性、有竞争力的身份保留能力、高效的流处理性能及稳定的长视频生成能力。

英文摘要

Talking-video character replacement requires coordinated transfer of appearance and voice while preserving the source motion, scene, linguistic content, and audio-video timing. Existing methods use separately optimized models for the two modalities, making audio-visual consistency difficult to enforce. We present UniSwap, the first framework for streaming joint audio-visual identity replacement in talking videos. Given a source video, a reference image, and a reference voice clip, UniSwap transfers the reference appearance and vocal timbre within a single audio-visual diffusion transformer while preserving the source content and dynamics. To address the scarcity of aligned cross-identity training pairs, we introduce a swap-and-reconstruct pipeline that removes visual and vocal identity from real clips and uses the original clips as reconstruction targets. Starting from a bidirectional backbone, we progressively adapt the model through In-context Pretraining for joint replacement, Conditional Streaming Adaptation for block-causal KV-cached generation, and Efficient Self-forcing DMD for mitigating exposure bias and reducing sampling from 30 to 3 denoising steps per block. Efficient Multi-LoRA Switching enables the three DMD roles to share a single frozen backbone. Feature-RoPE Decomposition keeps cached positions within the training range, supporting stable long-form inference. Experiments demonstrate strong audio-visual synchronization, competitive identity preservation, efficient streaming, and stable long-form generation.

发表机构

  • The Chinese University of Hong Kong(香港中文大学)
  • Qwen Applications Business Group of Alibaba(阿里巴巴通义千问应用业务集团)

机构由 AI 辅助整理,请以论文原文为准。

相关深度报道

↑