arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Logbook:极长时音频事件理解

Logbook: Extremely Long-form Audio Event Understanding

Kwanghee Choi, Suwon Shon, Dmitriy Serdyuk, Guitang Lan, Chao-Wei Huang, Mohammad Sadegh Rasooli, Sangeeta Srivastava, Zhaojiang Lin, Saurabh Adya, Ming Sun

arXiv 2610.07338首次发表:更新:

AI 中文总结

针对现有音频基准局限于短片段的问题,提出Logbook基准,涵盖十分钟至六天的录音,要求系统输出无间隙分割及事件标签和描述;通过比较52个系统,发现任务可行但最佳系统未达人类水平,且端到端优于级联,但长上下文下性能下降。

AI 中文摘要

音频基准测试通常围绕短时、预先分割的片段构建,这限制了模型设计只能处理简短输入或固定词汇表。为弥补这一差距,我们引入了Logbook,一个用于小时级音频理解的基准测试,其录音时长从十分钟到六天不等。给定一段连续音频录音和一个事件标签词汇表,系统必须预测一个无间隙的分割,并为每个片段提供事件标签和描述。我们比较了52个系统,包括端到端和级联系统,并消融了微调、上下文长度和推理预算。我们发现该任务是可处理的,尽管最佳系统仍低于人类参考水平。此外,过度分割现象普遍存在,微调能部分缓解这一问题。最后,端到端系统通常优于级联系统,但在更长的上下文下性能会下降。

英文摘要

Audio benchmarks are built around short, pre-segmented clips, limiting model design to brief inputs or fixed vocabularies. To close this gap, we introduce Logbook, a benchmark for hour-scale audio understanding, with recordings ranging from ten minutes to six days. Given a continuous audio recording and an event label vocabulary, a system must predict a gap-free segmentation with an event label and a description per segment. We compare 52 systems, end-to-end and cascaded, and ablate fine-tuning, context length, and reasoning budget. We find the task tractable, though the best systems remain below the human reference. Also, over-segmentation is pervasive, and fine-tuning partially mitigates it. Finally, end-to-end are often better than cascaded systems, but degrades with longer context.

CommentsSubmitted to ICASSP 2027. Source code available at https://github.com/facebookresearch/logbook

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑