AI 中文总结
该研究针对EuroLLM和Hy-MT2等5款1.7B至22B的翻译模型,结合文档分块与W4A8/W8A8量化,发现量化效率与翻译质量的权衡还受分块策略影响,且Hy-MT2比EuroLLM更耐量化。
AI 中文摘要
在实际服务器环境中部署大语言模型面临挑战,因为系统需提供高质量响应且延迟低。量化是减少内存占用、提升推理效率的常用方法,但在受控编排级工作负载下,其对延迟和吞吐量的影响很少被评估。本研究针对两个翻译模型系列——EuroLLM(Martins等,2025)和Hy-MT2(Zheng等,2026),涵盖1.7B至22B的5个模型,研究其在单块A100或H100 GPU上高效部署的量化权衡。研究表明,结合文档分块策略与W4A8或W8A8量化,可在广泛工作负载下优化延迟-吞吐量帕累托曲线。此外,由于标准机器翻译(MT)基准依赖孤立句子,无法捕捉长上下文动态,本研究引入WMT24++的文档级评估,以评估文本分块策略在量化下对翻译质量的影响。结果显示,标准句级评估无法预测量化与长文档翻译的相互作用:Hy-MT2在量化下保持鲁棒性,而EuroLLM对量化高度敏感,所有考虑的量化格式下翻译质量均快速下降。总体而言,推理效率与翻译质量的权衡不仅取决于量化格式,还取决于文本分块策略的选择。
英文摘要
Deploying large language models in realistic server environments poses challenges, as the system needs to provide high-quality responses with low latency. Quantization is a common approach to reduce the memory footprint and improve inference efficiency, yet its impact on latency and throughput is rarely evaluated under controlled, orchestration-level workloads. In this work we study the quantization trade-offs of EuroLLM \citep{martins2025eurollm} across three model sizes ranging from 1.7B to 22B for efficient deployment on a single A100 or H100 GPU. We demonstrate that combining a document-chunking strategy with W4A8 or W8A8 quantization improves the latency-throughput Pareto-curve under a wide range of workloads. Furthermore, since standard machine translation (MT) benchmarks rely on isolated sentences and fail to capture long-context dynamics, we introduce a document-level evaluation based on DocHPLT \cite{o2025dochplt} to assess how text chunking strategies affect translation quality under quantization. Our results indicate that standard segment-level evaluation can potentially underestimate the interaction between quantization and long-context document translation, for some quantization formats, translation direction and models. Overall, our experiments show that the trade-off between inference efficiency and translation quality depends not only on the quantization format, but also on the choice of text chunking strategy.