量化对大型语言模型孟加拉语理解能力的影响:一项系统评估
Quantization Effects on Bangla Language Understanding in Large Language Models: A Systematic Evaluation
浏览论文内容
中文总结 AI 辅助
本研究首次系统评估了三种量化格式对Qwen-2.5-7B等三个模型家族在五个孟加拉语NLU基准上的影响,发现架构和量化方法比位宽更关键,量化可用于孟加拉语部署。
中文摘要 AI 辅助
训练后量化可降低大型语言模型(LLMs)的内存占用并加快推理速度,因此如今常用于设备端部署。然而,我们对其影响的大部分认知来自英语基准测试。对于孟加拉语这种形态复杂、低资源的语言,同样的结论是否适用尚不明确,这正是本研究要解决的问题。我们通过lm-evaluation-harness工具,在零样本评估设置下,对三个模型家族——Qwen-2.5-7B、LLaMA-3.1-8B和GPT-OSS-20B——的全精度版本,以及三种量化格式(GPTQ-Int8、GPTQ-Q8、GGUF-W8A16)进行评估,覆盖五个孟加拉语自然语言理解基准(Bangla MMLU、CommonsenseQA-BN、OpenBookQA-BN、PIQA-BN和BoolQ-BN)。据我们所知,这是首次针对孟加拉语自然语言理解开展的量化格式控制对比研究。三个模型家族的表现存在差异:在GGUF-W8A16格式下,GPT-OSS在推理密集型任务上的准确率下降了57.35%,而Qwen和LLaMA在GPTQ格式下表现稳定,部分情况下量化版本的性能还略优于全精度版本;作为理解任务的BoolQ-BN,在所有模型家族和量化格式下均保持稳定。综合来看,这些结果表明量化技术可很好地应用于孟加拉语相关部署,但架构和量化方法的选择比单纯的位宽更为重要。我们还讨论了这一发现对在受限硬件上运行模型的从业者的意义。
英文摘要
Post-training quantization lowers the memory footprint of Large Language Models (LLMs) and speeds up inference, which is why it is now common for on-device deployment. Most of what we know about its effects, however, comes from English benchmarks. It is not clear whether the same holds for morphologically complex, low-resource languages such as Bangla, and this gap is what we address here. We evaluate three model families---Qwen-2.5-7B, LLaMA-3.1-8B, and GPT-OSS-20B---in full precision and in three quantized formats (GPTQ-Int8, GPTQ-Q8, GGUF-W8A16) across five Bangla natural language understanding benchmarks (Bangla MMLU, CommonsenseQA-BN, OpenBookQA-BN, PIQA-BN, and BoolQ-BN), using zero-shot evaluation through lm-evaluation-harness. To our knowledge this is the first controlled comparison of quantization formats on Bangla NLU. The three families do not respond the same way: GPT-OSS loses up to 57.35% accuracy on reasoning-heavy tasks under GGUF-W8A16, while Qwen and LLaMA hold steady under GPTQ, and in a few cases the quantized version edges out the full-precision one. BoolQ-BN, a comprehension task, stays stable across all three families regardless of format. Taken together, these results suggest quantization can work well for Bangla deployment, but the choice of architecture and quantization method matters more than the bit width alone. We discuss what this means for practitioners choosing a model to run on constrained hardware.
发表机构
- Institute of Information and Communication Technology (IICT)(信息与通信技术研究所)
- Shahjalal University of Science and Technology (SUST)(沙赫贾拉勒科技大学)
机构由 AI 辅助整理,请以论文原文为准。