发表机构
Meta Reality Labs(元宇宙现实实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对音频驱动面部动画的流式需求,提出FaceGAN模型,通过非因果噪声整形突破GAN的随机结构限制,实现单遍实时生成,性能优于或匹配现有最优方法且无漂移。
AI 中文摘要
音频驱动的面部动画是实时虚拟化身、远程呈现和具身虚拟智能体的核心,必须在线运行:每帧基于截至当前时间观测到的音频生成,且需达到交互速率。近期该领域进展以扩散模型为主,这类模型每个样本需要多次网络评估,因此不适合流式场景。我们认为该领域无需付出此成本:音频条件下的面部运动处于相对低维流形,单遍GAN即可满足需求,障碍并非容量而是随机结构。我们证明,由独立同分布噪声驱动的因果时不变生成器,无法在不破坏每步创新的情况下抑制输出频谱的某一频段。我们提出FaceGAN,通过对噪声通路进行非因果整形解决了该限制:由于驱动噪声是合成的,其未来值可提前采样,因此音频到表情的通路保持因果性,模型支持完全因果操作。FaceGAN每帧通过单次前向传播生成表情和头部姿态,在生成质量上与现有最优方法相当或更优;作为具有有限注意力窗口的前馈模型,它可无限生成且不会出现漂移。
英文摘要
Audio-driven facial animation underpins real-time avatars, telepresence, and embodied virtual agents. And it must run online: each frame emitted from audio observed up to the current time, at interactive rates. Recent progress is dominated by diffusion models, which need many network evaluations per sample and are therefore a poor fit for streaming. We argue the cost is unnecessary in this domain. Audio-conditioned facial motion occupies a comparatively low-dimensional manifold, a regime where a single-pass GAN suffices. The obstacle is not capacity but stochastic structure. We show that a causal, time-invariant generator driven by i.i.d. noise cannot suppress its output spectrum over a band without collapsing its per-step innovation. We proposed FaceGAN, which dissolved the limitation by shaping the noise pathway acausally. Because the driving noise is synthetic, its future can be sampled now, so the audio-to-expression path stays causal, and the model supports fully causal operation. FaceGAN emits expression and head pose in a single forward pass per frame and matches or outperforms state-of-art approaches in generation quality. Being feed-forward with bounded attention windows, it generates indefinitely without drift.
CommentsProject website: https://wojciechzielonka.com/facegan/