arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25890cs.CL

重新思考基于长度的训练:语音标记语言模型中的批次组成与损失归一化

Rethinking Length-Based Training: Batch Composition and Loss Normalization in Speech Token Language Models

Hongjin Song, Runwu Shi, Weiqiao Shan, Jiale Luo, Yujin Wang, Yifei Wu, Chunxiang Jin

首次发表
浏览论文内容

中文总结 AI 辅助

该研究通过匹配比较分离短到长训练中的批次组成、标记暴露和损失归一化因素,发现其独立收益有限,并提出系统分析协议。

中文摘要 AI 辅助

短到长训练是一种简单的语音模型课程策略,但其收益难以解释。在语音标记语言模型中,基于长度的训练在批次均值损失下可能改变洗牌策略、批次组成、标记保留和标记权重。我们通过匹配比较来分离这些因素。在测试设置中,当批次组成和标记暴露固定时,短到长排序没有显示出独立收益。在批次均值损失下,首轮分组降低了Mimi的困惑度,但在标记平衡损失下未观察到这一收益。跨分词器结果与块长度变化和标记权重之间的关联一致。这项工作为研究变长语音模型中基于长度的训练提供了系统分析协议。

英文摘要

Short-to-long training is a simple curriculum for speech models, but its gains can be difficult to interpret. In speech token language models, length-based training can change the shuffle policy, batch composition, token retention, and token weights under batch-mean loss. We disentangle these factors through matched comparisons. In the tested settings, short-to-long ordering shows no independent benefit when batch composition and token exposure are fixed. First-epoch grouping lowers perplexity for Mimi under batch-mean loss, but this gain is not observed under token-balanced loss. The cross-tokenizer results are consistent with a link between chunk-length variation and token weighting. This work provides a systematic analysis protocol for studying length-based training in variable-length speech models.

发表机构

  • Beijing Institute of Technology, Zhuhai(北京理工大学珠海校区)
  • Institute of Science Tokyo(东京科学大学)
  • Northeastern University(东北大学)
  • Sichuan University(四川大学)
  • Wuhan University(武汉大学)
  • Ant Group(蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

↑