arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

KanAdapter:一种基于柯尔莫哥洛夫-阿诺尔德网络的即插即用模块,用于高效微调基础语音模型

KanAdapter: A Kolmogorov-Arnold Network-based Plug-and-Play Module for Efficient Fine-tuning of Foundation Speech Models

Phuong Tuan Dat, Phuong Khai Minh, Tran Huy Dat

arXiv 2609.05281首次发表:更新:

发表机构

National University of Singapore; Institute of Advanced Intelligence and Computing (IAIC), A * STAR(新加坡国立大学; 新加坡科技研究局高级智能与计算研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

KanAdapter是基于GR-KAN的轻量级适配器,可高效微调语音基础模型,在多项任务中参数减少显著且性能优于MLP适配器,还能缓解持续学习的灾难性遗忘。

AI 中文摘要

对自监督学习(SSL)语音模型进行全量微调以完成下游任务的计算成本极高,而现有的参数高效微调方法主要依赖基于多层感知机(MLP)的适配器,其固定的激活函数在严格的参数预算下限制了表示表达能力。我们提出KanAdapter,这是一种轻量级适配器框架,用分组有理柯尔莫哥洛夫-阿诺尔德网络(GR-KAN)模块替代传统的MLP瓶颈,以实现更具表达性且参数效率更高的适配。遵循并行瓶颈设计,KanAdapter在冻结的Transformer编码器块旁插入可训练的GR-KAN分支,并利用预训练MLP层的权重实现稳定初始化。在说话人验证、语音情感识别和深度伪造检测任务中,KanAdapter相比全量微调可实现高达97.5%的可训练参数减少,同时保持高度竞争力,且在可比参数预算下始终优于AdaptFormer。在持续学习场景中,它相比全量微调及基于MLP的适配器可实现高达83.6%的误差降低,这归因于GR-KAN的有理激活的局部性,缓解了灾难性遗忘。据我们所知,这是首个探索基于KAN的模块用于语音基础模型参数高效微调的研究工作。

英文摘要

Fully fine-tuning self-supervised learning (SSL) speech models for downstream tasks is computationally prohibitive, and existing parameter-efficient fine-tuning approaches predominantly rely on MLP-based adapters whose fixed activation functions limit their representational expressiveness under tight parameter budgets. We propose \textbf{KanAdapter}, a lightweight adapter framework that replaces conventional MLP bottlenecks with Group-Rational Kolmogorov-Arnold Network (GR-KAN) modules for more expressive and parameter-efficient adaptation. Following a parallel bottleneck design, KanAdapter inserts trainable GR-KAN branches alongside frozen Transformer encoder blocks and leverages weight transfer from pre-trained MLP layers for stable initialization. Across speaker verification, speech emotion recognition, and deepfake detection, KanAdapter achieves up to 97.5\% reduction in trainable parameters relative to full fine-tuning while remaining highly competitive, and consistently outperforms AdaptFormer under comparable parameter budgets. In continual learning, it yields up to 83.6\% error reduction over full fine-tuning and MLP-based adapters, which we attribute to the localized nature of GR-KAN's rational activations that mitigates catastrophic forgetting. To our knowledge, this is the first work to explore KAN-based modules for parameter-efficient fine-tuning of speech foundation models.

CommentsAccepted to IMPACT-SPEECH Workshop @ EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑