arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2604.08384eess.AScs.AI

TASU2:用于语音大语言模型对齐和低资源适应的可控CTC模拟

TASU2: Controllable CTC Simulation for Alignment and Low-Resource Adaptation of Speech LLMs

  • X-LANCE Lab, Department of Computer Science and Engineering, Shanghai Jiao Tong University(上海交通大学计算机科学与工程系X-LANCE实验室)
  • MoE Key Lab of Artificial Intelligence(教育部人工智能重点实验室)
  • Jiangsu Key Lab of Language Computing(江苏省语言计算重点实验室)
  • AISpeech Ltd(思必驰科技股份有限公司)
  • Nanjing University(南京大学)

机构由 AI 辅助整理,请以论文原文为准。

Jing Peng, Chenghao Wang, Yi Yang, Lirong Qian, Junjie Li, Yu Xi, Shuai Wang, Kai Yu

更新

AI总结:

TASU2通过可控CTC模拟生成文本监督,提升语音大语言模型的对齐和低资源适应性能,优于TASU和文本细调等基线方法。

AI中文摘要:

语音大语言模型的后训练越来越依赖高效的跨模态对齐和稳健的低资源适应,但收集大规模音频文本对仍然成本高昂。文本仅对齐方法如TASU通过从转录中模拟CTC后验来减轻这一负担,但它们对不确定性和错误率的控制有限,使课程设计大多依赖启发式方法。我们提出了TASU2,一种可控的CTC模拟框架,能够在指定的WER范围内模拟CTC后验分布,生成文本衍生的监督,更好地匹配声学解码接口。这使得后训练课程能够平滑地变化监督难度而无需TTS。在多个源到目标适应设置中,TASU2在域内和域外识别上优于TASU,并且一致优于文本仅微调和基于TTS的增强等强基线方法,同时缓解了源域性能退化问题。

英文摘要:

Speech LLM post-training increasingly relies on efficient cross-modal alignment and robust low-resource adaptation, yet collecting large-scale audio-text pairs remains costly. Text-only alignment methods such as TASU reduce this burden by simulating CTC posteriors from transcripts, but they provide limited control over uncertainty and error rate, making curriculum design largely heuristic. We propose \textbf{TASU2}, a controllable CTC simulation framework that simulates CTC posterior distributions under a specified WER range, producing text-derived supervision that better matches the acoustic decoding interface. This enables principled post-training curricula that smoothly vary supervision difficulty without TTS. Across multiple source-to-target adaptation settings, TASU2 improves in-domain and out-of-domain recognition over TASU, and consistently outperforms strong baselines including text-only fine-tuning and TTS-based augmentation, while mitigating source-domain performance degradation.

↑