arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26431cs.SDeess.AS

AudioSpan:覆盖音频理解的时长与深度

LongAudioSpan: Spanning the Duration and Depth of Audio Comprehension

Wen Huang, Yunfei Chu, Meng Gao, Haolin He, Jin Xu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出覆盖10分钟至2小时音频的基准AudioSpan,含3240个分三认知层级的问题,评估12个大型音频-语言模型,发现从长音频提取相关事实是核心难点,该基准可公开获取。

中文摘要 AI 辅助

通用音频理解如今涵盖语音、声音和音乐,时长从几秒到几小时,这一进展由日益全模态的大型音频-语言模型(LALMs)推动。然而,测试这些模型的基准仍依赖几秒长的片段,导致分数饱和且模型收敛;近期的长音频研究虽延长了时长,但对长音频的评估方式与短片段类似。我们提出AudioSpan,一个同时覆盖时长与深度的基准:它将时长从10分钟到超过2小时的音频,与3240个分属感知、理解、推理三个认知层级的问题配对。问题通过两种路径生成,二者在问题内容来源和真实答案获取方式上存在差异:原生问答(Native QA)从音频内容中提取问题,每个问题以选择题和开放题形式呈现,开放题依据详细评分标准打分;锚定问答(Anchor QA)则注入真实答案,将声学锚点植入音频并构建感知到推理的链条,仅对首个错误进行评分。一条全自动化流水线通过结构化字幕生成、问答生成和对抗性评论反馈来构建每个条目。在AudioSpan上评估12个LALMs后,我们发现难点出现在推理之前:从冗长冗余的信号中提取少量相关事实。这一难点随音频时长增加而加剧,且在感知任务中最为突出,尤其是时间定位。AudioSpan可通过此httpsURL获取。

英文摘要

General audio comprehension now covers speech, sound, and music over durations from seconds to hours, driven by large audio-language models (LALMs) that are increasingly omni-modal. Yet the benchmarks that test them still rely on clips of seconds, where scores saturate and models converge; recent long-form efforts extend duration but evaluate long audio much as short clips are. We introduce LongAudioSpan, a benchmark that spans both duration and depth: it pairs audio from 10 minutes to over 2 hours with 3,240 questions across three cognitive levels, namely perception, understanding, and reasoning. Two paths supply the questions, differing in how question content is sourced and how ground truth is obtained. Native QA extracts questions from the audio's content, posing each as a multiple-choice item and an open-ended one graded by detailed rubrics. Anchor QA instead injects ground truth, planting acoustic anchors into the audio and building a perception-to-reasoning chain scored only to the first error. A fully automated pipeline constructs every item through structured captioning, QA generation, and adversarial critic feedback. Evaluating 12 LALMs on LongAudioSpan, we find the hard part comes before reasoning: distilling a few relevant facts from a long, redundant signal. This difficulty grows with audio length and falls hardest on perception, especially temporal grounding. LongAudioSpan is available at https://huggingface.co/datasets/holvan/LongAudioSpan.

发表机构

  • Qwen Team, Alibaba Group(通义千问团队,阿里巴巴集团)
  • Tsinghua University(清华大学)
  • The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

↑