AI 中文总结
针对音频时刻检索中DETR解码器缺乏全局归一化的问题,提出片段后验解码,通过前向-后向推理计算精确边际后验,在CASTELLA上显著提升性能。
AI 中文摘要
音频时刻检索(AMR)旨在从长录音中识别与自由形式文本查询最匹配的时间片段。现有系统主要依赖固定槽位的DETR解码器,这些解码器为提案分配置信度分数,但未对完整时间线的竞争性解释进行显式归一化。我们提出片段后验解码,该方法在时间片段划分上定义了一个全局归一化的分布,并通过前向-后向推理计算每个候选时刻的精确片段边际后验,以此对候选时刻进行评分。我们进一步通过将前景跨度及其相邻细分视为不同的假设来扩展训练片段空间,从而增加替代片段划分之间的竞争。在CASTELLA数据集上,我们的方法达到了41.15%的R1@0.7和34.68%的mAP,分别比使用DETR槽位置信度解码的同一网络高出10.91和9.20个百分点。
英文摘要
Audio moment retrieval (AMR) identifies temporal segments in long recordings that best match a free-form text query. Existing systems largely rely on fixed-slot DETR decoders that assign proposal-level confidence scores without explicitly normalizing over competing explanations of the full timeline. We propose segmental posterior decoding, which defines a globally normalized distribution over temporal segmentations and scores each candidate moment by its exact segment marginal posterior computed through forward-backward inference. We further expand the training segmentation space by treating a foreground span and its adjacent subdivisions as distinct hypotheses, thereby increasing competition among alternative segmentations. On CASTELLA, our method achieves 41.15% R1@0.7 and 34.68% mAP, outperforming the same network decoded with DETR slot confidence by 10.91 and 9.20 percentage points, respectively.
Comments5 pages, 2 figures, 2 tables. Submitted to ICASSP 2027