arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于LALM反馈的连续自回归非语言发声生成偏好优化

Preference Optimization with LALM Feedback for Continuous Autoregressive Non-Verbal Vocalization Generation

Jingbin Hu, Qirui Zhan, Yuang Cao, Ziyu Zhang, Yunxiang Chen, Houdun Liu, Su Feng, Bengu Wu, Lei Xie, Liumeng Xue

arXiv 2609.11260首次发表:更新:

发表机构

Northwestern Polytechnical University; Shenzhen Pimei Technology Co., Ltd.; Yutu Zhineng; Nanjing University(西北工业大学; 深圳湃美科技有限公司; 玉兔智能; 南京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出利用大型音频语言模型反馈进行偏好优化的框架,通过两阶段策略(RSFT和Anchored Flow-DPO)提升连续自回归模型中的非语言发声生成质量,在NVVSpeech挑战赛上超越基线。

AI 中文摘要

我们提出了一种基于大型音频语言模型(LALM)反馈的偏好优化框架,用于连续自回归语音模型中的可控非语言发声(NVV)生成。为构建无需人工偏好标注的偏好数据,我们结合注入NVV的真实转录文本与LLM生成的语义对齐提示,构建了双语提示语料库,执行随机模型展开,并使用LALM对候选话语进行排序,形成同提示下的选择-拒绝对。随后,我们采用两阶段优化策略:首先通过拒绝采样微调(RSFT)使模型适应LALM选择的高分样本,接着采用锚定流匹配直接偏好优化(Anchored Flow-DPO),利用话语级流匹配损失进行成对偏好优化,并保留选择样本的流匹配目标作为SFT锚点。该设计使得在无需显式序列似然的情况下实现DPO式偏好学习,同时保留对优选实现的直接监督。在官方1,600条话语的NVVSpeech挑战赛Track 2测试集上,我们的方法取得了最终Track2Score为75.80(中文79.39/英文72.21),比VoxCPM2基线高出+1.84。改进主要源于更高的NVV准确性和NVV感知效果,而整体质量保持稳定。

英文摘要

We propose a preference optimization framework with Large Audio-Language Model (LALM) feedback for controllable non-verbal vocalization (NVV) generation in continuous autoregressive speech models. To construct preference data without human preference annotation, we build a bilingual prompt corpus by combining NVV-injected real transcripts with LLM-generated semantically aligned prompts, perform stochastic model rollouts, and use a LALM to rank candidate utterances and form same-prompt chosen--rejected pairs. We then adopt a two-stage optimization strategy: Rejection Sampling Fine-Tuning (RSFT) first adapts the model to LALM-selected high-scoring samples, followed by Anchored Flow-DPO, which formulates pairwise preference optimization using utterance-level flow-matching loss and retains the chosen-sample flow-matching objective as an SFT anchor. This design enables DPO-style preference learning without explicit sequence likelihoods while preserving direct supervision on preferred realizations. On the official 1,600-utterance NVVSpeech Challenge Track~2 test set, our method achieves a Final Track2Score of \textbf{75.80} (79.39 ZH / 72.21 EN), outperforming the VoxCPM2 baseline by \textbf{+1.84}. The improvements are mainly driven by higher NVV Accuracy and NVV Perceptual Effect, while Overall Quality remains stable.

CommentsAccepted by ISCSLP 2026, NVVSpeech Challenge Track2

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑