arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于参考锚定的数据策展用于指令遵循泰英机器翻译

Reference-Grounded Data Curation for Instruction-Following Thai-English Machine Translation

Thodsaporn Chay-intr, Krittapad Harnchang, Mahannop Thabua, Kobkrit Viriyayudhakorn, Thanaruk Theeramunkong

arXiv 2609.34770首次发表:更新:

发表机构

iApp Technology; Intelligent Informatics and Service Innovation Research Center; Artificial Intelligence Entrepreneur Association of Thailand (AIEAT); Sirindhorn International Institute of Technology, Thammasat University(iApp科技公司; 智能信息学与服务创新研究中心; 泰国人工智能企业家协会; 诗琳通国际理工学院,法政大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对指令遵循泰英翻译中规则遵从与质量权衡问题,提出参考锚定数据策展两阶段流水线,构建Grounded数据集并微调出ChindaMT系列模型,在多个规模上优于或匹配基线,最高胜率68.4%。

AI 中文摘要

指令遵循机器翻译(IF-MT)要求遵守关于术语、格式和语域的提示级规则。规则遵从通常与翻译质量相互权衡,这种张力是通用目的的指令数据增强方法无法解决的。我们提出参考锚定数据策展(Reference-Grounded Data Curation),这是一个两阶段流水线,从已经满足监督约束的参考翻译中提取每条监督约束,从而通过构造确保可行性。第一阶段应用指令遵循难度(IFD)评分,从英泰平行语料池中保留最难但可学习的实例。第二阶段从每个参考目标中提取约束,并仅保留满足所有约束的生成结果,从而产生包含197万条记录的Grounded数据集。我们在Grounded上微调开放权重基础模型,产生ChindaMT,一个参数规模为4B、2B和0.8B的泰英翻译系列。在长度控制的成对评判下,ChindaMT在普通翻译和显式规则下均优于或匹配每个层级的每个同规模基线,对最强基线的胜率最高达68.4%。该方案可干净地迁移到Qwen各代模型。我们发布模型权重、Grounded数据集和评估套件。

英文摘要

Instruction-following machine translation (IF-MT) requires respecting prompt-level rules on terminology, formatting, and register. Rule compliance typically trades off against translation quality, a tension that general-purpose IF data augmentation methods do not address. We propose Reference-Grounded Data Curation, a two-phase pipeline that extracts every supervised constraint from a reference translation that already satisfies it, ensuring feasibility by construction. Phase 1 applies Instruction-Following Difficulty (IFD) scoring to retain the hardest-but-learnable instances from an English-Thai parallel pool. Phase 2 extracts constraints from each reference target and keeps only generations satisfying every constraint, yielding the 1.97M-record Grounded dataset. We fine-tune open-weight bases on Grounded to produce ChindaMT, a Thai-English translation family at 4B, 2B, and 0.8B parameters. Under length-controlled pairwise judging, ChindaMT outperforms or matches every same-size baseline at every tier on both plain translation and under explicit rules, reaching up to a 68.4% win rate against the strongest baseline. The recipe transfers cleanly across Qwen generations. We release model weights, the Grounded dataset, and evaluation suites.

CommentsAccepted at AACL-IJCNLP 2026 (Main Conference)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑