发表机构
LTCI, Télécom Paris, Institut Polytechnique de Paris(巴黎理工学院巴黎电信学院LTCI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出一个基准框架,系统评估大型音频-语言模型从原子声音事件组合推断更高级别人类活动的能力,发现当前模型无法仅凭音频可靠完成此类推理。
AI 中文摘要
大型音频-语言模型(LALMs)擅长对原子声音事件进行理解和推理任务,但它们从这些细粒度事件中推断更高级别人类活动的能力在很大程度上仍未得到检验。日常人类行为和活动,如摆桌子、打扫房屋或准备早餐,是从时间上分布的声音事件中组合涌现出来的,需要超越当前训练和评估范式所主导的事件中心粒度的抽象。我们的基准在一个原则性框架下评估了广泛的LALMs,该框架测试基于声学感知的语言推理如何将声音抽象结构化为更高级别的理解。通过系统地改变示例典型性和干扰物相似性,我们的评估揭示了当前模型仅从音频中无法可靠地从原子声学事件到更高级别人类活动执行组合推理。所有数据、分类法和评估脚本均可在我们的配套网站上公开获取:此 https URL
英文摘要
Large audio-language models (LALMs) excel at understanding and reasoning tasks over atomic sound events, yet their ability to infer higher-level human activities from such fine-grained events remains largely unexamined. Everyday human actions and activities, such as setting a table, cleaning the house, or preparing a breakfast emerge compositionally from temporally distributed sound events, requiring abstraction beyond the event-centric granularity that dominates current training and evaluation paradigms. Our benchmark evaluates a wide set of LALMs under a principled framework that tests how language-based reasoning, grounded in acoustic perception, structures sound abstractions into higher-level understanding. By systematically varying exemplar typicality and distractor similarity, our evaluation exposes \added{that current models do not reliably perform compositional inference from atomic acoustic events to higher-level human activities solely from audio.} All data, taxonomies, and evaluation scripts are publicly available on our companion website: https://alm-sounding-actions.onrender.com/
Journal refThe 2026 Conference on Empirical Methods in Natural Language Processing, Oct 2026, Budapest, Hungary