arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高效地使口语语言模型适应新加坡语境

Efficiently Adapting Spoken Language Models for the Singaporean Context

Ng Jia Sheng Jason

arXiv 2607.10092首次发表:更新:

发表机构

xData Home Team Science & Technology Agency (HTX)(新加坡内政团队科技局)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对原始训练数据不可用且需多语言口语交互的情况,将开源口语语言模型适应新加坡语境。核心方法是结合LoRA微调等技术,构建多语言数据集。贡献是HT - Moonstone(5B)在多数任务表现出色,口音和性别识别佳,且原始语音问答能力损失小。

AI 中文摘要

口语语言模型统一了语音感知和推理,但将其应用于敏感领域的研究较少,特别是在原始训练数据无法获取且用例需要多语言口语查询交互的情况下。我们将一个开源口语语言模型适应新加坡内政团队的语境,涉及新加坡四种官方语言的五项语音任务,结合了LoRA微调、防止灾难性遗忘的替代文本问答数据集以及使CoBa重加权方案适应语音的多任务目标。我们还构建了HTD - 多语言 - QA,一个包含504,853个样本的文本和口语形式的多语言问答数据集。结果得到的HT - Moonstone(5B)在大多数任务上匹配或超越了高达其7倍规模的口语语言模型,在所有评估模型中获得了最佳的口音和性别识别能力,并且仅损失了不到2%的原始语音问答能力。

英文摘要

Spoken language models (SLMs) unify speech perception and reasoning, but adapting them to sensitive domains is underexplored, especially when the original training data is inaccessible and the use case demands multilingual, spoken-query interaction. We adapt an open-source SLM to the Singaporean Home Team context across five speech tasks in Singapore's four official languages, combining LoRA fine-tuning, a surrogate text-QA dataset that guards against catastrophic forgetting, and a multi-task objective that adapts the CoBa reweighting scheme to speech. We also build HTD-multilingual-QA, a 504,853 sample multilingual QA dataset in text and spoken form. The resulting HT-Moonstone (5B) matches or outperforms SLMs up to 7x its size on most tasks, attains the best accent and gender recognition among all models evaluated, and loses under 2\% of its original speech QA ability.

Comments10 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑