发表机构
Tencent Video; The University of Hong Kong(腾讯视频; 香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对联合音视频扩散模型在多奖励强化学习中固定路由和权重无法适应训练动态的问题,提出自适应奖励路由,通过跨模态影响引导定位更新和偏好保留重加权协调奖励,显著提升模态质量、语义一致性与音视频同步。
AI 中文摘要
多奖励引导的强化学习(即RL)为沿着多个互补目标改进联合音视频扩散模型提供了一种有前景的方法,这些目标包括模态特定质量、跨模态语义对齐和时间同步。然而,其有效性取决于训练过程中变化的两个量:奖励驱动的更新应在何处作用,以及竞争性奖励应如何组合。现有方法往往依赖固定的路由和奖励权重,无法跟踪不断演化的模型功能。为解决这些局限,我们提出自适应奖励路由,在前向过程强化学习(即DiffusionNFT)中联合调整联合音视频扩散模型的更新位置和奖励协调。我们的方法包含两个组件。(i)跨模态影响引导路由(定位更新):我们使用双向交叉注意力响应作为演化中跨模态影响的高效代理,动态重新加权token级损失并在跨模态层间缩放梯度,无需额外模型干预。(ii)保留偏好的模态感知重加权(协调奖励):我们将预定义权重保留为偏好先验,并在预热后使用分支特定的奖励梯度交互作为残差校正。这解决了演化中的冲突,而不让主导奖励抑制弱但必要的目标。大量实验证明,在模态质量、语义一致性和音视频同步方面,相对于强RL基线有一致改进。消融研究和机制分析进一步验证了自适应更新路由和奖励协调的互补优势。
英文摘要
Multi-reward guided reinforcement learning (i.e., RL) offers a promising way to improve joint audio-video diffusion models along several complementary objectives, including modality-specific quality, cross-modal semantic alignment, and temporal synchronization. Its effectiveness, however, depends on two quantities that change during training: where reward-driven updates should act, and how competing rewards should be combined. Existing methods tend to rely on fixed routing and reward weights, failing to track evolving model functions. To address these limitations, we propose Adaptive Reward Routing to jointly adapt update locations and reward coordination during forward-process RL (i.e., DiffusionNFT) of joint audio-video diffusion models. Our method consists of two components. (i) Cross-Modal Influence-Guided Routing (Localizing Updates): We use bidirectional cross-attention responses as an efficient proxy for evolving cross-modal influence, dynamically reweighting token-aware losses and scaling gradients across cross-modal layers without additional model interventions. (ii) Preference-Preserving Modality-Aware Reweighting (Coordinating Rewards): We preserve predefined weights as preference priors and use branch-specific reward-gradient interactions as residual corrections after warm-up. This resolves evolving conflicts without letting dominant rewards suppress weak but essential objectives. Extensive experiments demonstrate consistent improvements in modality quality, semantic consistency, and audio-video synchronization over strong RL baselines. Ablations and mechanism analyses further validate the complementary benefits of adaptive update routing and reward coordination.