arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

直接偏好密度对齐用于对话音频均衡

Direct Preference Density Alignment for Conversational Audio Equalization

Ioannis Stylianou, Sven Ewan Shepstone, Jon Francombe, Pablo Martinez Nuevo, Zheng-Hua Tan

arXiv 2609.12607首次发表:更新:

发表机构

Aalborg University; Bang & Olufsen A/S(奥尔堡大学; Bang & Olufsen公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出直接偏好密度对齐框架,无需代理奖励模型,结合GRPO在线探索与DPO离线细化,在音频均衡任务中以1.5B参数模型达到GPT-4o mini的感知水平。

AI 中文摘要

大语言模型对齐通常依赖于学习得到的代理奖励模型,这会显著增加训练期间的内存占用,并且极易出现不稳定性和奖励黑客问题。虽然像直接偏好优化(DPO)这样的离线方法绕过了奖励模型,但它们失去了进行在线探索的能力。如果不施加优化约束,这可能导致在有界连续空间中发生格式崩溃。为解决此问题,我们提出了直接偏好密度对齐:一种替代框架,它在严格保留在线强化学习优势的同时,消除了对学习得到的代理奖励模型的需求。我们利用大规模用户数据(约90,000个样本)构建非参数偏好密度图,从而建立一个经验奖励曲面。除了移除奖励模型外,直接偏好密度对齐还使得能够将群体相对策略优化(GRPO)的在线结构基础与DPO的针对性离线细化相结合。我们表明,这种GRPO+DPO组合达到了最高性能,并且在盲听音频均衡测试中,使一个1.5B参数的模型仅使用一小部分推理计算量,就能在感知上与精心设计提示的GPT-4o mini基线达到同等水平。

英文摘要

Large Language Model alignment typically relies on learned proxy reward models, which significantly increase the memory footprint during training and are notoriously prone to instability and reward hacking. While offline methods like Direct Preference Optimization (DPO) bypass the reward model, they lose the ability to perform online exploration. If no optimization constraints are applied, this can lead to format collapse in bounded, continuous spaces. To resolve this, we propose Direct Preference Density Alignment: An alternative framework that removes the need for a learned proxy reward model while strictly preserving the benefits of online reinforcement learning. We leverage large-scale user data (approximately 90,000 samples) to construct non-parametric preference density maps, establishing an empirical reward surface. In addition to removing the reward model, Direct Preference Density Alignment enables the combination of the online structural grounding of Group Relative Policy Optimization (GRPO) with the targeted offline refinement of DPO. We show that this GRPO+DPO combination achieves the highest performance, and in a blind audio equalization listening test, enables a 1.5B-parameter model to achieve perceptual parity with a carefully prompt-engineered GPT-4o mini baseline, using only a fraction of the inference compute.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑