arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02831cs.SDcs.CL

以演化 rubric 作为奖励的强化学习用于音频推理

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning

Fangxu Yu, Tao Feng, Dehai Min, Zinan Lin, Weijia Xu, Michael Xu, Philip S. Yu, Ge Liu, Tianyi Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

提出 AudioRubrics 强化学习框架,以自演化的音频基 rubric 奖励监督音频推理,在三个基准上大幅优于基线,收敛至稳定推理长度,提升音频感知效果。

中文摘要 AI 辅助

音频推理是机器理解声学世界的关键。带有可验证奖励的强化学习可引导此类推理,但现有奖励设计存在互补性局限:基于结果的奖励仅监督最终答案,使模型无需关注音频即可得到答案;基于过程的奖励对推理本身评分,但依赖粗糙、手工设计且固定的标准,这些标准既不适应每个问题,也不基于声学证据。此外,不同问题需求各异,部分依赖感知,部分依赖多步推理,任何静态标准都会随策略改进而失效。因此,用细粒度、基于音频且自适应的奖励监督推理过程至关重要,但手动为每个样本设计此类奖励不切实际,颇具挑战。为此,我们提出 AudioRubrics,这是一个用自演化、基于音频的 rubric 奖励监督音频推理的强化学习框架。AudioRubrics 从原始波形合成每个样本的 rubric,并根据模型自身的 rollout(回滚)为每组重新生成和重新加权标准,提供持续的学习信号,在静态标准饱和时持续针对当前策略的弱点。对三个音频推理基准的综合评估显示,AudioRubrics 大幅优于大量开源及基于训练的基线。此外,我们的分析表明,性能提升随 rubric 生成器和判别器的能力而扩展,且 AudioRubrics 收敛至稳定的推理长度,避免了退化崩溃和无限制增长。音频感知的改进进一步证明了将监督锚定在声学证据中的有效性。我们的项目页面可通过此 https URL 访问。

英文摘要

Audio reasoning is essential for machine understanding of the acoustic world. Reinforcement learning with verifiable rewards can elicit such reasoning, yet existing reward designs are complementary in their limitations: outcome-based rewards supervise only the final answer and let the model reach it without attending to the audio, whereas process-based rewards score the reasoning itself but rely on coarse, hand-crafted, and fixed criteria that neither adapt to each question nor stay grounded in the acoustic evidence. Moreover, questions differ in what they demand, with some hinging on perception and others on multi-step reasoning, and any static criterion weakens as the policy improves. Supervising the reasoning process with fine-grained, audio-grounded, and adaptive rewards is therefore crucial, yet challenging since such rewards are impractical to design by hand for every sample. To this end, we introduce AudioRubrics, a reinforcement learning framework that supervises audio reasoning with self-evolving, audio-grounded rubric rewards. AudioRubrics synthesizes per-sample rubrics from the raw waveform and, conditioned on the model's own rollouts, regenerates and reweights criteria per group, supplying a continuous learning signal that keeps targeting the current policy's weaknesses as static criteria saturate. Comprehensive evaluations across three audio reasoning benchmarks reveal that AudioRubrics substantially outperforms a wide range of open-source and training-based baselines. Furthermore, our analysis shows that the gains scale with the capability of the rubric generator and judge, and AudioRubrics converges to a stable reasoning length that avoids both degenerate collapse and unbounded growth. The improvement in audio perception further demonstrates the effectiveness of anchoring supervision in the acoustic evidence. Our project page is available at https://audiorubrics.github.io.

发表机构

  • University of Maryland(马里兰大学)
  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • University of Illinois Chicago(伊利诺伊大学芝加哥分校)
  • Microsoft Research(微软研究院)
  • MBZUAI(穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

相关深度报道

↑