arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

引导语音语言模型:通过对比激活添加实现免训练任务特化

Steering Speech-Language Models: Training-Free Task Specialization via Contrastive Activation Addition

Séverin Baroudi, Yanis Labrak, Pierfrancesco Melucci, Sergio Burdisso, Petr Motlicek, Hervé Bredin, Mirco Ravanelli, Ricard Marxer

arXiv 2610.04683首次发表:更新:

发表机构

Univ Toulon, Aix Marseille Univ, LIS, CNRS; Mila – Quebec AI Institute; Idiap Research Institute; Sapienza University of Rome; Brno University of Technology; pyannoteAI; Concordia University; CNRS, ILLS(土伦大学,艾克斯-马赛大学,LIS,法国国家科学研究中心; Mila – 魁北克人工智能研究所; Idiap 研究所; 罗马大学; 布尔诺理工大学; pyannoteAI 公司; 康考迪亚大学; 法国国家科学研究中心,ILLS)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出免训练的对比激活添加(CAA)协议,从少量带标签语音样本中为SpeechLLMs推导引导向量,在推理时增强目标语音任务,与提示结合可提升ASR、ER等任务性能,并支持跨域迁移和脚本归一化。

AI 中文摘要

激活引导已被证明在推理时能有效控制大型语言模型(LLMs)的行为,但其在语音语言模型(SpeechLLMs)中的应用仍属新兴领域,针对此类模型的免训练引导方法在很大程度上尚未被探索。我们提出了一种免训练的对比激活添加(CAA)协议,该协议从少量带标签的语音样本中为SpeechLLMs中的常见语音任务(如转录)推导出引导向量。我们展示了在推理时将这些向量添加到表示空间中,能更好地强制执行目标语音任务。我们进一步表明,当与提示(prompting)结合时,这些向量在大多数评估任务(如自动语音识别(ASR)或情感识别(ER))上相比单独使用提示带来了一致的改进,并能迁移到域外数据。我们还展示了脚本归一化方向在强制特定语言的目标脚本方面的实用性。

英文摘要

Activation steering has proven effective for controlling the behavior of Large Language Models (LLMs) at inference time, but its application to SpeechLLMs remains new, and training-free steering approaches for such models are still largely unexplored. We propose a training-free Contrastive Activation Addition (CAA) protocol that derives steering vectors for common speech tasks (e.g. transcription) in SpeechLLMs from a small number of labeled utterances. We showcase that adding these vectors in the representation space, at inference time, enforces better the targeted speech task. We further show that, when combined with prompting, these vectors yield to consistent improvement over prompting alone on most evaluated tasks such as Automatic Speech Recognition (ASR) or Emotion Recognition (ER), and transfer to out-of-domain data. We additionally demonstrate the usefulness of script-normalization directions to enforce the target script of a specific language.

CommentsSubmitted to ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑