TQLite:多大语言模型评审团引导的蒸馏用于实时MQM翻译质量评估
TQLite: Multi-LLM Jury Guided Distillation for Real-time MQM Translation Quality Evaluation
浏览论文内容
中文总结 AI 辅助
本研究提出TQLite蒸馏框架,通过多LRM评审团生成合成训练数据,使SLMs的MQM翻译质量评估性能远超普通SLM,成为LLM/LRM评估器的高性价比替代方案。
中文摘要 AI 辅助
大型语言模型(LLMs)在基于MQM的翻译质量(TQ)评估中展现出令人印象深刻的性能,近期大型推理模型(LRMs)的进展有望带来更大提升。然而,LLMs和LRMs大规模部署的计算成本高昂,而小型语言模型(SLMs)虽效率高,却难以完成评估任务所需的复杂推理。本研究开展了广泛的实证研究,在多种TQ评估设置中对SLMs、LLMs和LRMs进行基准测试,全面呈现当前研究现状并确立最佳实践。为解决可扩展性挑战,我们提出TQLite这一新型蒸馏框架,使SLMs能够达到基于最优LRM的评估器的MQM评估性能。该方法利用多LRM评审团,通过实用的数据整理技术及多样化模型群体的评估响应聚合生成高质量合成训练数据。结果表明,经TQLite训练的SLMs具备强大的MQM评估性能,远超标准SLM的现成评估能力,为LLM和LRM评估器提供了可扩展且高性价比的替代方案。
英文摘要
Large language models (LLMs) have demonstrated impressive performance in MQM-based translation quality (TQ) evaluation, and recent advances in large reasoning models (LRMs) promise even greater improvements. However, both LLMs and LRMs are computationally expensive to deploy at scale, while small language models (SLMs)---though much more efficient---struggle with the complex reasoning required for evaluation tasks. In this work, we present an extensive empirical study benchmarking SLMs, LLMs, and LRMs across a wide range of TQ evaluation setups, providing a comprehensive view of the current landscape and establishing best practices. To address the scalability challenge, we introduce TQLite, a novel distillation framework that enables SLMs to approach the MQM evaluation performance of the best LRM-based evaluators. Our approach leverages a multi-LRM jury to generate high-quality synthetic training data via practical data curation techniques and aggregation of evaluation responses across a diverse panel of models. Our results demonstrate that SLMs trained via TQLite achieve strong MQM evaluation performance that far exceeds off-the-shelf evaluation capabilities of standard SLMs, offering a scalable and cost-effective alternative to LLM- and LRM-based evaluators.
发表机构
- Netflix(网飞公司)
机构由 AI 辅助整理,请以论文原文为准。