当指标奖励最差翻译:内化文化推理用于社交媒体翻译评估
When Metrics Reward the Worst Translations: Internalizing Cultural Reasoning for Social Media Translation Evaluation
浏览论文内容
中文总结 AI 辅助
针对通用翻译指标在社交媒体文化负载内容上失效的问题,提出CuRIL强化学习框架,通过内化文化推理训练评判模型,以8B参数接近Pro级模型性能,并将低质量翻译率降低超20个百分点。
中文摘要 AI 辅助
在通用领域语料上训练出的自动翻译质量指标,系统性地在社交媒体内容上失效,因为社交媒体的交际意图编码在文化负载表达(网络俚语、谐音密码和平台特定习语)中,而非表面词元模式。我们进行了系统的实证分析,证明包括COMET、XCOMET和BERTScore在内的标准指标与人类文化判断呈现接近零或负相关,甚至出现严重性反转,即分数随翻译质量下降而升高。我们进一步表明,这种失败也延伸至大语言模型评判者:Qwen3-235B的Cohen's kappa仅为0.162,揭示瓶颈不是推理能力而是文化根基:模型缺乏识别翻译中哪些方面需要审查所需的领域特定文化知识。为解决此问题,我们提出CuRIL,一种内化文化推理的强化学习框架:文化注释被前置到模型推理内部,通过词元级损失掩码从策略梯度中排除,并以随训练衰减至零的概率注入,逐步强制自主文化判断。在包含1,444个样本的人工标注社交媒体翻译基准上,使用CuRIL训练的Qwen3-8B达到Cohen's kappa 0.370和精确匹配准确率45.22%,以30倍更少的参数接近Gemini-3.1-Pro,并超越规模高达235B的模型。我们进一步证明,我们的评判者为下游翻译优化提供可靠奖励信号,在独立人工评估下将低质量翻译率降低超过20个百分点。
英文摘要
Automatic translation quality metrics trained on general-domain corpora systematically fail on social media content, where communicative intent is encoded in culturally loaded expressions (internet slang, homophonic ciphers, and platform-specific idioms) rather than surface token patterns. We conduct a systematic empirical analysis demonstrating that standard metrics including COMET, XCOMET, and BERTScore exhibit near-zero or negative correlation with human cultural judgments, and even display a severity inversion in which scores increase as translation quality deteriorates. We further show that this failure extends to large language model judges: Qwen3-235B achieves Cohen's kappa of only 0.162, revealing that the bottleneck is not reasoning capacity but cultural grounding: models lack the domain-specific cultural knowledge needed to identify which aspects of a translation require scrutiny. To address this, we propose CuRIL, a reinforcement learning framework that internalizes cultural reasoning: cultural annotations are prepended inside the model's reasoning, excluded from policy gradients via a token-level loss mask, and injected with a probability that decays to zero over training, progressively forcing autonomous cultural judgment. On a 1,444-sample human-annotated social media translation benchmark, Qwen3-8B trained with CuRIL achieves Cohen's kappa 0.370 and Exact Match accuracy of 45.22%, approaching Gemini-3.1-Pro with 30x fewer parameters and surpassing models up to 235B in scale. We further demonstrate that our judge produces reliable reward signals for downstream translation optimization, reducing the low-quality translation rate by over 20 percentage points under independent human evaluation.
发表机构
- Zhejiang University(浙江大学)
- Xiaohongshu Inc.(小红书科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。