arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新审视音频时刻检索中的帧级显著性

Revisiting Frame-Wise Saliency for Audio Moment Retrieval

Tatsuya Komatsu, Hokuto Munakata

arXiv 2610.05737首次发表:更新:

发表机构

LY Corporation(LY公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文证明帧级显著性序列可直接用于音频时刻检索,在CASTELLA数据集上优于解码器输出,尤其对短时刻提升显著,且解码器监督仍有助于训练。

AI 中文摘要

本文重新审视了音频时刻检索(AMR)中的帧级显著性。我们表明,帧级显著性序列在基于DETR的AMR模型中通常仅用作辅助输出,但其本身可以作为时刻预测的有效来源。我们使用一种简单的受SED启发的分割规则(无学习参数)将显著性序列转换为排序时刻,从而直接从帧级时间信息实现时刻检索。在CASTELLA数据集上,基于显著性的预测在QD-DETR和CG-DETR的所有18次运行中始终优于来自同一模型的基于解码器的预测。对于QD-DETR,仅替换推理输出即可将R1@0.7从21.0提高到36.1。当两种输出在其独立选择的最佳epoch进行评估时,优势仍保持7-15个百分点。同样的趋势也扩展到TaskWeave和UVCOM,而TR-DETR则表现出相反的行为,这表明显著性的构建方式可能很重要。性能差距在短时刻中尤为明显:对于注释时刻平均最多2秒的查询,使用QD-DETR时R1@0.7从6.3提高到27.8。然而,解码器监督仍然有利于基于显著性的预测,表明其在训练期间的作用与其推理输出的效用不同。

英文摘要

This paper revisits frame-wise saliency for audio moment retrieval (AMR). We show that the frame-wise saliency sequence, conventionally used only as an auxiliary output in DETR-based AMR models, can itself serve as an effective source of moment predictions. We convert the saliency sequence into ranked moments using a simple SED-inspired segmentation rule with no learned parameters, enabling moment retrieval directly from frame-wise temporal information. On the CASTELLA dataset, saliency-based prediction consistently outperforms decoder-based prediction from the same model across all 18 runs of QD-DETR and CG-DETR. For QD-DETR, simply replacing the inference output improves R1@0.7 from 21.0 to 36.1. The advantage remains 7-15 points when the two outputs are evaluated at their independently selected best epochs. The same tendency extends to TaskWeave and UVCOM, whereas TR-DETR shows the opposite behavior, suggesting that how saliency construction may matter. The performance gap is especially pronounced for short moments: for queries whose annotated moments average at most 2 s, R1@0.7 improves from 6.3 to 27.8 with QD-DETR. Decoder supervision nevertheless benefits saliency-based prediction, indicating that its role during training differs from the utility of its inference output.

CommentsICASSP2027 submission

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑