arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.09681cs.SD

超越准确率:用于评估大型音频语言模型音频推理能力的ARIA-Rubrics

Beyond Accuracy: ARIA-Rubrics for Evaluating Audio Reasoning in Large Audio Language Models

  • Imperial College London(帝国理工学院)
  • Technische Universität München(慕尼黑工业大学)
  • Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
  • Shanghai Jiao Tong University(上海交通大学)
  • Johns Hopkins University(约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

Yupei Li, Qiyang Sun, Mohamed Mady, Chenxi Wang, Zhengwei Gong, Berrak Sisman, Björn Schller

中文总结 AI 辅助

针对大型音频语言模型推理评估难题,提出ARIA-Rubrics框架,利用六个指标和思维链提示,实现无需标注的自动透明评估,并识别出三种推理模式。

中文摘要 AI 辅助

大型音频语言模型(LALMs)在音频推理基准测试中表现出色,但仅凭准确率无法区分真正的推理与表面的模式匹配,且往往高估推理能力,因为高分可能源于猜测而非真正的音频理解。评估推理过程本身对于提升LALMs的推理能力至关重要,但仍具挑战性。现有方法要么依赖昂贵的人工标注,要么采用不透明的LLM-as-judge方法,这些方法不实用、有偏见且缺乏透明度。此外,音频推理引入了文本环境中不存在的独特挑战,即感知幻觉以及音频理解与文本推理之间的跨模态对齐,因此基于文本的评估框架无法直接应用。为此,我们提出ARIA-Rubrics(音频推理完整性评估),一个轻量级、无需标注、基于黄金推理链的自动且透明的框架,包含六个互补指标,从感知基础、推理连贯性和答案一致性三个维度评估音频推理质量。我们使用思维链提示作为外部化机制,使推理过程可观察。在2个基准上的9个模型进行的实验识别出当前LALMs的三种推理模式,并为未来发展提供了可操作的方向,ARIA-Rubrics与人类判断具有高度相关性。代码可在GitHub仓库获取。

英文摘要

Large Audio Language Models (LALMs) have shown strong performance on audio reasoning benchmarks, but accuracy alone cannot distinguish true reasoning from superficial pattern matching, often overestimating reasoning ability since high scores may result from guessing rather than genuine audio understanding. Evaluating the reasoning process itself is essential for improving LALMs' reasoning ability, yet remains challenging. Existing methods either rely on costly human annotation or opaque LLM-as-judge approaches, making them impractical, biased, and lacking transparency. Moreover, audio reasoning introduces unique challenges absent in text-based settings, perceptual hallucination and cross-modal alignment between audio understanding and textual inference, hence text-based evaluation frameworks cannot be directly applied. Therefore, we propose ARIA-Rubrics (Audio Reasoning Integrity Assessment), a lightweight, annotation-free gold reasoning chains, automatic and transparent framework comprising six complementary metrics that evaluate audio reasoning quality across perceptual grounding, reasoning coherence, and answer consistency. We use Chain-of-Thought prompting as an externalization mechanism to make the reasoning process observable. Experiments on 9 models across 2 benchmarks identify three reasoning modes of current LALMs with actionable directions for future development, with ARIA-Rubrics achieving high correlation with human judgments. The code is available at the Github Repository.

补充信息

↑