发表机构
South East Technological University; IRISA(东南理工大学; IRISA研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对生物医学领域法语到英语翻译,联合应用知识蒸馏与量化压缩模型,实现体积减少69%、推理速度提升98.21%、碳排放减少98.46%且不损失翻译质量。
AI 中文摘要
大规模预训练Transformer模型已在多种机器翻译任务中取得了最先进的性能,包括多语言场景。知识蒸馏已成为一种可持续的模型压缩方法,将知识从大型教师模型转移到更小、更高效的学生模型。类似地,量化通过降低模型权重和激活值的数值精度(例如,从32位表示降至8位表示)被广泛用于加速推理,使模型在部署时运行速度提升数倍。然而,这两种技术在应用于专业领域数据时均面临局限性,尤其是在低资源条件下。在知识蒸馏中,迁移的有效性往往受到领域特定平行数据稀缺的制约,而量化则可能随着位精度降低而导致性能下降。在本工作中,我们研究了知识蒸馏与量化在法语到英语生物医学翻译中的联合应用,该领域以专业术语和有限的平行资源为特征。我们开发并比较了多种微调策略,以使压缩后的学生模型适应这一具有挑战性的场景。我们的实验表明,协同蒸馏和量化的学生模型相比原始基线,模型体积减少了69%,推理速度提升了98.21%,二氧化碳排放量减少了98.46%,且翻译质量未受影响。这些结果表明,联合优化的压缩技术能够产生高效、高性能的模型,适用于在资源受限条件下运营的翻译服务提供商。
英文摘要
Large-scale pretrained transformer models have achieved state-of-the-art performance across diverse machine translation tasks, including multilingual settings. Knowledge distillation has emerged as a sustainable approach for model compression, transferring knowledge from large teacher models to smaller, more efficient student models. Similarly, quantization, which reduces the numerical precision of model weights and activations (e.g., from 32-bit to 8-bit representations) is widely used to accelerate inference, enabling models to run several times faster during deployment. However, both techniques face limitations when applied to specialized domain data, particularly under low-resource conditions. In knowledge distillation, the effectiveness of transfer is often constrained by the scarcity of domain-specific parallel data, while quantization can lead to performance degradation as bit precision decreases. In this work, we investigate the combined application of knowledge distillation and quantization for French-to-English biomedical translation, a domain characterized by specialized terminology and limited parallel resources. We develop and compare multiple fine-tuning strategies to adapt compressed student models to this challenging setting. Our experiments demonstrate that a collaboratively distilled and quantized student model achieves a 69% reduction in size, a 98.21% increase in inference speed, and a 98.46% reduction in CO2 emissions compared to the original baseline all without sacrificing translation quality. These results indicate that jointly optimized compression techniques can yield efficient, high-performance models suitable for translation service providers operating under resource constraints.
CommentsAccepted at AICS 2025