arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AudioMap:用于时间感知密集音频字幕生成的完形填空与选择强化学习方法

AudioMap: Cloze-and-Choice Reinforcement Learning for Time-Aware Dense Audio Captioning

Yan Rong, Fengji Ma, Xu Li, Jinting Wang, Chen Zhang, Li Liu

arXiv 2608.09559首次发表:更新:

AI 中文总结

针对时间感知密集音频字幕生成的现有方法缺陷,提出基于完形填空与选择奖励范式的AudioMap框架,构建首个相关数据集AudioMapCap-44K,在基准测试中取得最优性能。

AI 中文摘要

时间感知密集音频字幕生成(TDAC)旨在生成音频的多个细粒度属性(密集)并附带精确的时间边界(时间感知)。现有方法难以同时实现这两个目标,且主要依赖监督微调,导致性能次优。虽然强化学习(RL)具有应用潜力,但将其应用于TDAC面临两大挑战:(1)现有奖励过于粗糙,无法以细粒度方式监督多事件、多属性和多关系描述;(2)自由形式字幕的时间监督存在困难,灵活的事件时间表达式使得可靠的事件时间对应关系难以建立。为解决这些挑战,我们提出AudioMap,一种基于RL的新型TDAC框架,其采用统一的完形填空与选择奖励范式。具体而言,我们引入证据充分性奖励(ESR),采用非对称分层评分机制,以提升不同声学维度下的细粒度准确性和描述丰富度。此外,我们设计了事件条件时间奖励(ECTR),通过时间交并比(IoU)将时间戳与事件语义进行结构化绑定,并辅以双课程学习策略以促进训练过程。最后,为支持该任务,我们构建了首个时间感知细粒度音频字幕数据集AudioMapCap-44K,包含44K条精心标注的字幕。在不同基准上的大量实验表明,AudioMap在开源模型中达到了最先进(SOTA)性能,且相对于专有模型取得了具有竞争力或更优的结果。项目页面和发布更新可在该https URL获取。

英文摘要

Time-aware dense audio captioning (TDAC) aims to generate multiple fine-grained attributes (dense) of the audio with precise time boundaries (time-aware). Existing methods struggle to achieve these two goals and mainly rely on supervised fine-tuning, yielding sub-optimal performance. While reinforcement learning (RL) shows promise, applying it to TDAC faces two main challenges: (1) existing rewards are too coarse to supervise multi-event, multi-attribute, and multi-relation descriptions in a fine-grained manner; and (2) temporal supervision is difficult for free-form captions, where flexible event-time expressions make reliable event-time correspondence challenging. To address these challenges, we propose AudioMap, a novel RL-based TDAC framework, which shifts to a unified cloze-and-choice reward paradigm. Specifically, we introduce the Evidence Sufficiency Reward (ESR) with an asymmetric hierarchical scoring mechanism to promote fine-grained accuracy and descriptive richness across diverse acoustic dimensions. Furthermore, we design the Event-Conditioned Temporal Reward (ECTR) to structurally bind timestamps to event semantics via temporal IoU, accompanied by a dual-curriculum learning strategy to facilitate the training process. Finally, to support this task, we construct the first time-aware fine-grained audio captioning dataset, AudioMapCap-44K, which contains 44K carefully annotated captions. Extensive experiments across diverse benchmarks show that AudioMap achieves state-of-the-art (SOTA) performance among open-source models and delivers competitive or superior results relative to proprietary models. Project page and release updates are available at https://github.com/ryysayhi/AudioMap.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑