发表机构
UX Factory, Inc.(UX工厂有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文跨架构评估文本转语音后训练量化,发现相同位宽因模型敏感组件不同而结果迥异,提出分阶段消融识别并逐层GPTQ修复,强调需在目标运行时验证。
AI 中文摘要
后训练量化(PTQ)降低了设备端文本转语音(TTS)的成本,但已发表的评估仅覆盖单一系统或方法。我们在一个统一协议下评估了跨TTS架构的PTQ,涉及三个核心模型、对另外八个模型进行权重和激活消融,以及两个盲量化的保留模型。四比特逐通道权重使Supertonic的UTMOS(一种预测平均意见得分)降低了2.8,而Kokoro降低了0.07;逐张量缩放即使在8比特下也可能导致严重退化。相同位宽产生不同结果,因为敏感组件是模型特定的,且无法从模型类别可靠预测。分阶段消融程序可识别该组件,逐层GPTQ可将其恢复至UTMOS 0.1以内。真实的int8和int4内核在硬件相关成本下复现了模拟排序。在Mac mini上,4比特权重内核运行Supertonic的速度为fp32延迟的0.60倍,而int8更慢,因此每个配置都需要在目标运行时上进行验证。
英文摘要
Post-training quantization (PTQ) reduces the cost of on-device text-to-speech (TTS), but published evaluations cover one system or method. We evaluate PTQ across TTS architectures under one protocol with three core models, weight and activation ablations of eight more, and two held-out models quantized blind. Four-bit per-channel weights reduce UTMOS, a predicted mean opinion score, by 2.8 on Supertonic and 0.07 on Kokoro, and per-tensor scaling can cause severe degradation even at 8 bits. The same bit width yields different outcomes, because the sensitive component is model-specific and not reliably predicted from the model class. A staged ablation procedure identifies it, and per-layer GPTQ can restore it to within 0.1 UTMOS. Real int8 and int4 kernels reproduce the simulated ordering at hardware-dependent cost. On a Mac mini, a 4-bit weight kernel runs Supertonic at 0.60x the fp32 latency while int8 is slower, so each configuration requires validation on the target runtime.
CommentsSubmitted to ICASSP 2027. 4 pages plus references. Code and run records: https://github.com/uxfacdev/tts-ptq-map