arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分层声学-语义建模:全双工口语语言模型的模态分离与语义连贯

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs

Zhenyu Liu, Xuanyu Zhang, Yunxin Li, Qixun Teng, Shenyuan Jiang, Haolan Chen, Minjun Zhao, Fanbo Meng, Yu Xu, Yancheng He, Baotian Hu, Haizhou Li, Min Zhang

arXiv 2607.06540首次发表:更新:

发表机构

School of Computer Science and Technology, Harbin Institute of Technology, Shenzhen; Center for Language, Intelligence and Machines, Shenzhen Loop Area Institute, Shenzhen; School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen(哈尔滨工业大学(深圳)计算机科学与技术学院; 深圳环智中语言、智能与机器中心; 香港中文大学(深圳)人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究全双工口语语言模型受模态干扰问题,提出Lychee-FD框架及分层参数分离策略,通过实验验证该方法能有效提升模型性能,在语音智能和交互流畅性上取得显著进步,且不影响推理效率。

AI 中文摘要

开发无缝、高性能、原生智能的全双工口语语言模型(SLMs)一直是语音和自然语言处理社区的关键挑战和长期目标。尽管有显著进展,但近期努力受严重模态干扰根本限制,导致知识退化和语义完整性受损。本文通过对模型优化动态的详尽细粒度分析,揭示性能下降根源是声学和语义建模在共享深度参数空间时的固有梯度冲突。基于此,引入Lychee-FD框架,提出分层参数分离策略,在多个全双工基准测试上实验表明该方法显著提升了技术水平,在语音智能和全双工交互流畅性上有大幅改进且不影响推理效率。

英文摘要

Developing seamless, high-performance, native intelligent full-duplex Spoken Language Models (SLMs) remains a critical challenge and long-standing goal for the speech and NLP community. Despite notable progress, recent endeavors are fundamentally constrained by severe modality interference, which causes substantial knowledge degradation and compromises semantic integrity -- ultimately making full-duplex SLMs feel unnatural and unintelligent. In this paper, through an exhaustive fine-grained analysis of model optimization dynamics, we uncover the root cause of such performance degradation, revealing that modality interference arises from inherent gradient conflicts between acoustic and semantic modeling when the two modalities are forced to share a deep parameter space. Guided by this key insight, we introduce Lychee-FD, a native end-to-end full-duplex framework designed to mitigate modality interference. Importantly, we propose a hierarchical parameter separation strategy that decouples conflicting modalities in deep layers while preserving cross-modality coherence via a dedicated semantic alignment channel. Extensive experiments on multiple full-duplex benchmarks demonstrate that our method significantly advances the state of the art, yielding substantial improvements in both speech intelligence (+7.4% on Spoken QA) and full-duplex interaction fluidity (+28.5% on FullDuplexBench 1.5) without compromising inference efficiency. To the best of our knowledge, this work is the first to achieve two key advances: 1) uncovering and elucidating the root cause of modality interference in full-duplex SLMs, and 2) designing an elegant hierarchical model together with a practical solution for seamless, high-performance, native intelligent full-duplex SLMs.

Comments22 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑