arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11913cs.CV

HarmoniDPO:基于偏好优化扩散模型的视频引导音频生成

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion

  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Wenshuo Peng, Kaipeng Zhang

AI总结:

本文提出HarmoniDPO框架,采用双视频表示、online-DPO微调及DDS算法,实现更精准的视频到音频生成,在音视频同步性与主观质量上优于现有最优方法。

AI中文摘要:

视频到音频(V2A)生成面临两大核心挑战:一是视觉与听觉线索间存在复杂模糊的关联,难以实现精准的时序同步;二是生成音频的感知质量难以达标。现有方法通常将视频输入压缩为单一特征表示,导致时序动态性与细粒度视觉信息大量丢失,且依赖基于重构的训练目标,与人类对音频质量及适配性的感知判断相关性差。针对这些局限,本文提出HarmoniDPO框架,将基于偏好的优化融入基于扩散的V2A生成任务中:(1)采用双视频表示,结合全局上下文与逐帧特征,以保留时序动态性和语义细节;(2)受人类反馈强化学习(RLHF)启发,HarmoniDPO采用在线直接偏好优化(online-DPO),基于偏好判断对扩散型V2A模型进行微调,提升音频的感知质量与对齐度;(3)此外,本文还引入双尺度扩散搜索(DDS)作为测试时缩放算法,可在推理阶段自适应优化输出保真度。实验结果表明,HarmoniDPO在音视频同步性与主观音频质量上均优于现有最优方法,为从视频生成符合人类偏好的真实音频提供了鲁棒解决方案。

英文摘要:

Video-to-audio (V2A) generation faces significant challenges in achieving precise temporal synchronization and high perceptual quality due to the complex, ambiguous relationship between visual and auditory cues. Existing methods typically compress video inputs into single feature representations, leading to significant loss of temporal dynamics and fine-grained visual information. These approaches also rely on reconstruction-based training objectives that poorly correlate with human perceptual judgments of audio quality and appropriateness. We propose HarmoniDPO, a novel framework that integrates preference-based optimization into diffusion-based V2A generation to address these limitations. (1) Our approach leverages a dual video representation: combining global context with frame-wise features to preserve temporal dynamics and semantic detail. (2) Inspired by reinforcement learning from human feedback (RLHF), HarmoniDPO employs online Direct Preference Optimization (online-DPO) to fine-tune a diffusion-based V2A model from preference judgments, enhancing perceptual quality and alignment. (3) Additionally, we introduce Dual-scale Diffusion Search (DDS), a test time scaling algorithm that adaptively optimizes output fidelity during inference. Experiments demonstrate that HarmoniDPO outperforms state-of-the-art methods in audio-video synchronization and subjective audio quality, offering a robust solution for generating realistic, human-preferred audio from video.

补充信息

↑