FlexiSLM: A Spoken Language Model with Dynamic and Controllable Frame Rates
FlexiSLM:一种动态可控帧率的语音语言模型
机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) ; ByteDance(字节跳动)
AI总结 提出FlexiSLM,首个支持动态可控帧率的语音语言模型,利用动态帧率表示在高质量点超越固定帧率7B模型,并在6.25 Hz时推理速度减半且保持高质量。
Comments Accepted to EMNLP2026 Main Conference