arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Alpha作为效率信号:可见性路由的RGBA图像到视频生成

Alpha as an Efficiency Signal: Visibility-Routed RGBA Image-to-Video Generation

Zhe Li, Honghao Qiao, Zhixin Xu, Qijie Wang, Bo Peng, Dawei Li

arXiv 2608.09355首次发表:更新:

发表机构

ByteDance(字节跳动)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对游戏RGBA视频生成的数据集与效率问题,提出可见性路由的RGBA视频生成模型,实现了更低FVD与1.2倍推理加速,质量下降可忽略。

AI 中文摘要

RGBA视频将RGB外观与Alpha通道结合,使动画资产可应用于任意背景,在游戏行业中被大量使用。然而,为游戏生成高质量的RGBA动画仍面临两大挑战:其一,多数现有RGBA视频数据集以逼真内容为主,对游戏资产的覆盖有限;其二,传统的“生成后遮罩”流水线仅在RGB合成后估算Alpha,导致半透明区域常被背景模糊,产生不稳定的遮罩输出。近年来,许多方法开始联合建模RGB与Alpha,但现有方法多以文本为条件,且在效率和质量上仍存在未解决的问题。为应对这些挑战,我们引入GameAlpha-2.4K,这是一个包含2400个片段的游戏风格RGBA视频数据集,采用了利于遮罩的合成、多假设Alpha恢复以及基于合成的质量控制机制。基于该数据集,我们训练了一个参考条件RGBA视频生成器,可单次生成RGB帧与Alpha遮罩。为提升效率,我们提出可见性路由(visibility router),在早期阶段识别透明令牌并跳过后续的DiT更新,同时x₀-lock引导其沿原始流匹配调度向自预测端点推进。我们的模型获得了比传统两阶段流水线更低的FVD,且可见性路由在最后两个DiT去噪步骤中跳过了35%的令牌评估,与密集推理相比,实现了1.2倍的主干网络加速,且质量下降可忽略不计。

英文摘要

RGBA videos combine RGB appearance with an alpha channel, enabling animated assets to be applied across arbitrary backgrounds, which are heavily used in gaming industry. However, generating high-quality RGBA animations for games remains challenging for two reasons. First, most existing RGBA video datasets are dominated by photorealistic content, with limited coverage of game assets. Second, the traditional generate-then-matte pipelines estimate alpha only after RGB synthesis, so semi-transparent regions are often blurred by background, resulting in unstable matting outputs. More recently, many methods have begun to model RGB and alpha jointly, but existing approaches are mostly text-conditioned, and still have unresolved issues in efficiency and quality. To address these challenges, we introduce GameAlpha-2.4K, a 2.4K-clip game-style RGBA video dataset built with matte-friendly synthesis, multi-hypothesis alpha recovery, and compositing-based quality gates. Using this dataset, we train a reference-conditioned RGBA video generator that jointly produces RGB frames and alpha mattes in a single pass. To improve efficiency, we propose a visibility router that identifies transparent tokens in an early stage and bypasses their later DiT updates, while x_0-lock guides them along the original flow-matching schedule toward self-predicted endpoints. Our model obtains lower FVD than traditional two-stage pipelines, and the visibility router skips 35% of token evaluations in the final two DiT denoising steps, providing a 1.2x backbone speedup with negligible quality degradation compared to dense inference.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑