arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ParA-LLM:副语言与声学语音理解的统一方法

ParA-LLM: A Unified Approach to Paralinguistic and Acoustic Speech Understanding

Nishit Anand, Jiaqi Su, Ke Chen, Yunyun Wang, Dinesh Manocha, Ramani Duraiswami, Rithesh Kumar, Zeyu Jin

arXiv 2609.22771首次发表:更新:

发表机构

Adobe Research; University of Maryland, College Park; OpenAI(Adobe研究院; 马里兰大学学院公园分校; OpenAI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对音频大模型忽视副语言信息的问题,提出ParA-LLM,通过两阶段课程训练和120万问答对,在ParA-Bench等基准上超越GPT-4o-Audio,显著提升说话人与声学特征理解。

AI 中文摘要

近期音频大语言模型(Audio LLMs)的进展已实现人类水平的语音识别,然而现有系统难以捕捉副语言方面,如说话人特质、表达变化及环境声学条件。为解决此问题,我们设计了一个包含22种副语言特征的框架,并创建了超过120万条音频问答对的数据集。我们开发了ParA-LLM,采用两阶段课程训练:首先在单属性问题上构建基础知识,然后在多属性问题上进行说话人与声学特征的联合推理。我们还发布了ParA-Bench基准,包含6,000道多选题,涵盖说话人语音、声学及混合类别,其中前沿模型如GPT-4o-Audio仅达到36%的准确率。ParA-LLM在ParA-Bench上超越最先进的音频大语言模型如GPT-4o-Audio达7.5%,并在MMAU-Pro Speech上额外提升1.13%,在MMAR Speech上提升7.49%。

英文摘要

Recent advances in Audio LLMs have achieved human-level speech recognition, yet existing systems struggle to capture paralinguistic aspects such as speaker traits, expressive variations, and environmental acoustic conditions. To address this, we design a framework of 22 paralinguistic characteristics and create a dataset of over 1.2M Audio-QA pairs. We develop ParA-LLM, trained with a two-stage curriculum: first on single-attribute questions to build foundational knowledge, then on multi-attribute questions for joint reasoning over speaker and acoustic characteristics. We also release ParA-Bench, a benchmark of 6,000 multiple-choice questions across speaker-speech, acoustic, and mixed categories, where frontier models like GPT-4o-Audio achieve only 36% accuracy. ParA-LLM surpasses state-of-the-art Audio LLMs like GPT-4o-Audio by 7.5% on ParA-Bench, with additional gains of 1.13% on MMAU-Pro Speech and 7.49% on MMAR Speech.

CommentsAccepted to Interspeech 2026. Project Website: https://nishitanand.github.io/paralinguistic-understanding-llm/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑