发表机构
Tongji University; The Chinese University of Hong Kong, Shenzhen; Peking University(同济大学; 香港中文大学(深圳); 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出HyPASE双曲PEFT框架,利用庞加莱球模型及HGA、EMCA组件,在MELD、IEMOCAP等数据集上优于欧氏PEFT基线,实现高效的大音频语言模型语音情感微调。
AI 中文摘要
大音频语言模型(Large Audio-Language Models, LALMs)在通用语音理解方面表现出色,但将其适配到语音情感识别(Speech Emotion Recognition, SER)这类细粒度任务仍是重大瓶颈。当前参数高效微调(Parameter-Efficient Fine-Tuning, PEFT)方法通常在平坦欧氏空间中运行,该几何空间无法捕捉从低层韵律到高层语义的多粒度情感线索特性。为解决此问题,本文提出HyPASE——一种基于LALM的SER双曲PEFT框架。HyPASE利用庞加莱球模型,将双曲半径作为表征粒度的显式代理。该框架包含两个核心组件:用于层自适应权重调制的双曲几何适配器(Hyperbolic Geometric Adapter, HGA),以及将多尺度特征压缩为紧凑音频前缀的情感感知多容量跨模态聚合器(Emotion-aware Multi-capacity Cross-modal Aggregator, EMCA)。在标准基准上的实验结果显示,HyPASE在MELD数据集的所有指标上均优于欧氏PEFT基线,且在IEMOCAP数据集上取得显著的未加权准确率提升,尤其在类别不平衡的情感识别任务中,伴随的轻微加权准确率折衷反映了双曲空间对少数类表征的几何优先级;此外,HyPASE在受限参数预算内实现了稳健的零样本跨数据集泛化。通过将微调过程建立在双曲几何基础上,HyPASE为LALM微调提供了一种高效路径。
英文摘要
Large Audio-Language Models (LALMs) excel at general speech understanding; however, adapting them to fine-grained tasks like Speech Emotion Recognition (SER) remains a significant bottleneck. Current Parameter-Efficient Fine-Tuning (PEFT) methods typically operate in flat Euclidean space, and this geometry fails to capture the multi-granularity nature of emotion cues, which range from low-level prosody to high-level semantics. To address this, we propose HyPASE, a hyperbolic PEFT framework for LALM-based SER. HyPASE leverages the Poincare ball model, using the hyperbolic radius as an explicit proxy for representational granularity. The framework consists of two core components: a Hyperbolic Geometric Adapter (HGA) for layer-adaptive weight modulation, and an Emotion-aware Multi-capacity Cross-modal Aggregator (EMCA) that compresses multi-scale features into compact audio prefixes. Empirical results on standard benchmarks show that HyPASE outperforms Euclidean PEFT baselines across all metrics on MELD and achieves a notable Unweighted Accuracy gain on IEMOCAP, particularly in class-imbalanced emotion recognition, with the accompanying slight Weighted Accuracy trade-off reflecting hyperbolic space's geometric prioritization of minority-class representations; furthermore, HyPASE achieves robust zero-shot cross-dataset generalization within a constrained parameter budget. By grounding the adaptation process in hyperbolic geometry, HyPASE offers a highly efficient path for LALM fine-tuning.