arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19932cs.CLcs.SD

通过渐进压缩实现口语语言模型的高效模态链推理

Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models

Pengchao Feng, Chao-Hong Tan, Qian Chen, Wen Wang, Xiangang Li, Xie Chen

首次发表
浏览论文内容

中文总结 AI 辅助

针对口语语言模型推理能力落后问题,提出高效模态链推理(ECoM推理),通过压缩文本组件提高推理准确性,并用渐进压缩策略训练,实验显示其在口语数学问答基准测试中,增强推理且保持效率,准确率提升显著。

中文摘要 AI 辅助

口语语言模型(SLMs)实现了自然的人机交互,但其推理能力仍落后于基于文本的大语言模型,特别是在口语数学问答任务上。原因在于SLMs对纯语言化数学表达式进行推理,比符号文本更难解释。由于架构限制和额外计算需求,直接将基于文本的推理转移到SLMs并不容易。为此提出高效模态链推理(ECoM推理),首个将压缩推理引入SLMs的框架。通过压缩文本组件,使其同时作为语音指导和推理表示,提高了推理准确性,且使用的令牌预算比标准模态链(CoM)架构小。还提出渐进压缩策略来训练该能力。实验表明,ECoM推理在无显式推理时比标准CoM提高了21%的准确率,在有完整推理痕迹时比CoM提高了3%,同时仅使用40%的文本令牌,证明其在增强SLM推理的同时保持推理效率。

英文摘要

Spoken language models (SLMs) enable natural human-computer interaction, but their reasoning ability still lags behind that of text-based large language models, especially on spoken mathematical question answering tasks. One important reason is that SLMs reason over purely verbalized mathematical expressions, which are harder to interpret than symbolic text. However, directly transferring text-based reasoning to SLMs is nontrivial due to architectural constraints and the additional computational requirements. To address this challenge, we propose Efficient Chain-of-Modality Reasoning (ECoM Reasoning), the first framework to introduce compressed reasoning into SLMs. By compressing the textual component so that it jointly serves as speech guidance and reasoning representation, ECoM Reasoning improves reasoning accuracy while using a smaller token budget than the standard Chain-of-Modality (CoM) architecture, which generates intermediate text before speech. To train this capability, we further propose Progressive Compression, a curriculum-based strategy that gradually trains the model from full-form reasoning to compressed reasoning. Experiments on spoken mathematical question answering benchmarks show that ECoM Reasoning improves accuracy by 21% over standard CoM without explicit reasoning, and by 3% over CoM with full reasoning traces while using only 40% of the text tokens, demonstrating that it enhances SLM reasoning while remaining inference-efficient.

发表机构

  • Shanghai Jiao Tong University(上海交通大学)
  • Shanghai Innovation Institute(上海创新研究院)
  • Token Foundry, Alibaba Group(阿里巴巴集团淘系技术)

机构由 AI 辅助整理,请以论文原文为准。

↑