arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12310cs.CLcs.AIcs.LG

ESTS 在 WMT26:基于路由信息的专家剪枝用于模型压缩

ESTS at WMT26: Routing-Informed Expert Pruning for Model Compression

  • University of California, Los Angeles(加州大学洛杉矶分校)

机构由 AI 辅助整理,请以论文原文为准。

Liu O. Martin, Lucas Bandarkar, Nanyun Peng

AI总结:

本文提出基于路由信息的专家剪枝方法,在WMT26模型压缩任务中压缩GPT-OSS-20B,通过任务路由质量排序和跨语言路由差异分配容量,结合恢复微调与MXFP4量化,实现参数4.186B至7.770B的高效压缩模型。

AI中文摘要:

我们描述了团队名称为 ESTS 的六项提交,参与无约束的 WMT26 模型压缩共享任务,涉及英语到简体中文和英语到埃及阿拉伯语两个翻译方向。我们为每个翻译方向提交了三个压缩操作点,全部基于 GPT-OSS-20B 模型。我们使用任务特定的路由质量对专家进行排序,并利用跨语言路由差异在层间分配保留容量,然后物理移除低重要性的专家。由此产生的专家模型在 GPT-5.1 生成的合成翻译数据上进行恢复微调,并通过将 MXFP4 量化应用于保留的专家投影权重进一步压缩。我们还为指令条件下的 WMT26 设置实现了一个稳健的推理系统,包括类别推断、输出验证、重试、分段回退和源端拥有的 JSON 重建。在我们的六项提交中,参数数量从 4.186B 到 7.770B 不等,打包工件大小从 4.55 到 6.33 GiB。使用 GPT-5.1 伪参考的内部 xCOMET-XL 评估提供了所提交压缩操作点之间的内部比较。

英文摘要:

We describe six submissions under the team name ESTS to the unconstrained WMT26 Model Compression Shared Task for English--Simplified Chinese and English--Egyptian Arabic. We submit three compression operating points per translation direction, all derived from GPT-OSS-20B. We use task-specific routing mass to rank experts and cross-lingual routing divergence to allocate retained capacity across layers, then physically remove low-importance experts. The resulting specialists are recovery-tuned on GPT-5.1-generated synthetic translation data and further compressed by applying MXFP4 quantization to the retained expert projection weights. We additionally implement a robust inference system for the instruction-conditioned WMT26 setting, including category inference, output validation, retries, segmented fallback, and source-owned JSON reconstruction. Across our six submissions, parameter counts range from 4.186B to 7.770B and packed artifact sizes from 4.55 to 6.33~GiB. Internal xCOMET-XL evaluation using GPT-5.1 pseudo-references provides an internal comparison across the submitted compression operating points.

补充信息

↑