arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26823cs.SDcs.CLeess.AS

文本分数可能遗漏波形使用:Qwen2-Audio 量化案例研究

Text Scores Do Not Establish Performance on Lexically Non-Diagnostic Speech Tasks: A Qwen2-Audio Quantization Case Study

  • National Research Council Canada(加拿大国家研究委员会)
  • The Hong Kong Polytechnic University(香港理工大学)
  • The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

Mengzhe Geng, Jinxi Ji, Junhao Xu

AI总结:

本研究通过 Qwen2-Audio 案例,提出评估协议,揭示文本分数可能遗漏波形依赖行为,表明所选量化分配在词汇任务上提升但情感识别上受损。

AI中文摘要:

语音语言模型的训练后量化通常以文本输出分数和名义位宽来概括。仅凭这些数字并不能确定依赖于转录中缺失信息的行为,也不能确定特定运行时的效率。我们引入了一种评估协议,分别测试词汇输出、转录不足的端点和实测的打包实现。在 Qwen2-Audio 案例研究中,一个翻译选择的 6 位分配在冻结的英译德回放上将 chrF 提高了 2.36,配对 95% 自助区间为 [1.04, 3.62],但在说话人分离的情感识别上损失了 3.91 个百分点。在相同的 6 位预算下,均匀结构控制达到的情感准确率高于所选分配,且前层控制在相同冻结集上的点估计也更高。在 7 位时,chrF 提高了 3.28,区间为 [2.08, 4.59],情感区间相对于 FP16 包含零,且同预算前层控制仍超过所选分配。一项单独的匹配预算 4.08 位研究发现,每个测试的低位分配都有大约 10 个点的情感缺陷,且所选分配相对于冻结控制没有优势。最后,一个反量化的平均 6 位模拟保留了 FP16 的峰值内存。本案例研究确定了词汇输出、波形依赖行为和名义精度之间依赖于精度的不匹配。它并未确立低位语音模型的普遍失败,也未确立所选分配的部署优势。

英文摘要:

Text-output scores alone do not show whether quantization preserves performance on speech tasks whose target labels cannot be recovered from the transcript. We evaluate fixed mixed 4/8-bit Qwen2-Audio-7B-Instruct allocations averaging 6 and 7 bits per parameter on 508 English-to-German FLEURS utterances and on 512 RAVDESS emotion clips from 16 speakers. The BLEU and chrF differences from half precision (FP16) have intervals that include zero for both allocations. On RAVDESS, the same two sentences occur equally often with every emotion label. The absolute accuracy differences from FP16 are -3.71% for 6 bit and -1.17% for 7 bit. The 6-bit speaker interval excludes zero and an exact two-sided sign-flip test gives p=0.0148; the 7-bit interval includes zero. Same-budget controls do not identify either selected allocation as best. This case study shows why translation scores and performance on tasks beyond the transcript need separate evaluation.

↑