arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AV-GRPO:用于联合音视频生成的模态锚定解耦扩散强化学习

AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation

Zhiyu Xu, Weilong Yan, Yufei Shi, Shiyang Li, Yihao Liu, Kin-Man Lam, Yuewen Cao

arXiv 2609.29816首次发表:更新:

发表机构

Shanghai AI Laboratory; The Hong Kong Polytechnic University; National University of Singapore; Nanyang Technological University; Zhejiang University(上海人工智能实验室; 香港理工大学; 新加坡国立大学; 南洋理工大学; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对联合音视频生成中模态保真度、文本对齐和同步不足的问题,提出模态锚定解耦扩散强化学习框架AV-GRPO及解耦数据集5DAV,在JavisBench和VABench上超越LTX-2.3。

AI 中文摘要

近年来,联合音视频生成领域取得了重大进展。现有模型仍然存在模态保真度有限、文本模态对齐不足以及跨模态同步较弱的问题。虽然强化学习后训练提供了一种有前景的补救措施,但直接将其应用于联合音视频生成具有挑战性。异构多模态奖励会纠缠学习信号并使信用分配复杂化。鉴于两种模态塔的动态差异,联合优化它们在计算上代价高昂。此外,同步评估的难度取决于配对样本,阻碍了公平的奖励比较。我们提出了AV-GRPO,一种模态锚定的在线扩散强化学习框架,以及5DAV,一个解耦的、难度可控的训练数据集。AV-GRPO包含三个关键模块:(1)模态锚定的rollout,用于解缠学习信号并稳定难度;(2)轨迹锁定的冻结塔优化,以降低成本并重新分配信用;(3)针对模态特定动态定制的自适应目标和扰动强度。这将耦合的多模态偏好学习转化为单模态子问题,以实现精确的奖励归因和更好的同步。我们的5DAV数据集在五个维度上解耦样本,用于系统训练。在JavisBench和VABench上的实验表明,在LoRA和全微调下,AV-GRPO在生成质量、语义对齐和跨模态同步方面优于LTX-2.3。消融实验证实了我们的设计。代码和数据:此https URL

英文摘要

Recent years have witnessed major progress in joint audio-video generation. Existing models still suffer from limited per-modality fidelity, insufficient text-modality alignment and weak cross-modal synchronization. While reinforcement-learning post-training offers a promising remedy, directly adapting it to joint audio-video generation is challenging. Heterogeneous multimodal rewards entangle learning signals and complicate credit assignment. Joint optimization of two modality towers is computationally expensive given their divergent dynamics. Moreover, synchronization evaluation difficulty depends on paired samples, preventing fair reward comparisons. We propose AV-GRPO, a modality-anchored online diffusion RL framework, and 5DAV, a decoupled, difficulty-controllable training dataset. AV-GRPO includes three key modules: (1) modality-anchored rollouts to disentangle learning signals and stabilize difficulty; (2) trajectory-locked frozen-tower optimization to reduce cost and reassign credit; (3) adaptive objectives and perturbation strengths tailored to modality-specific dynamics. This converts coupled multimodal preference learning into unimodal subproblems for precise reward attribution and better synchronization. Our 5DAV dataset decouples samples across five dimensions for systematic training. Experiments on JavisBench and VABench demonstrate AV-GRPO outperforms LTX-2.3 in generation quality, semantic alignment and cross-modal synchronization under LoRA and full fine-tuning. Ablations confirm our designs. Code and data: https://github.com/zhiyuxu03/AV-GRPO

Comments22 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑