arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08156cs.CL

当指标奖励最差翻译:内化文化推理用于社交媒体翻译评估

When Metrics Reward the Worst Translations: Internalizing Cultural Reasoning for Social Media Translation Evaluation

Yiwen Qiu, Linjuan Wu, Dingming Li, Yizhou Liu, Zixuan Wang, Haolei Xu, Ye Guo, Daoxin Zhang, Weiming Lu, Yongliang Shen

首次发表
浏览论文内容

中文总结 AI 辅助

针对通用翻译指标在社交媒体文化负载内容上失效的问题,提出CuRIL强化学习框架,通过内化文化推理训练评判模型,以8B参数接近Pro级模型性能,并将低质量翻译率降低超20个百分点。

中文摘要 AI 辅助

在通用领域语料上训练出的自动翻译质量指标,系统性地在社交媒体内容上失效,因为社交媒体的交际意图编码在文化负载表达(网络俚语、谐音密码和平台特定习语)中,而非表面词元模式。我们进行了系统的实证分析,证明包括COMET、XCOMET和BERTScore在内的标准指标与人类文化判断呈现接近零或负相关,甚至出现严重性反转,即分数随翻译质量下降而升高。我们进一步表明,这种失败也延伸至大语言模型评判者:Qwen3-235B的Cohen's kappa仅为0.162,揭示瓶颈不是推理能力而是文化根基:模型缺乏识别翻译中哪些方面需要审查所需的领域特定文化知识。为解决此问题,我们提出CuRIL,一种内化文化推理的强化学习框架:文化注释被前置到模型推理内部,通过词元级损失掩码从策略梯度中排除,并以随训练衰减至零的概率注入,逐步强制自主文化判断。在包含1,444个样本的人工标注社交媒体翻译基准上,使用CuRIL训练的Qwen3-8B达到Cohen's kappa 0.370和精确匹配准确率45.22%,以30倍更少的参数接近Gemini-3.1-Pro,并超越规模高达235B的模型。我们进一步证明,我们的评判者为下游翻译优化提供可靠奖励信号,在独立人工评估下将低质量翻译率降低超过20个百分点。

英文摘要

Automatic translation quality metrics trained on general-domain corpora systematically fail on social media content, where communicative intent is encoded in culturally loaded expressions (internet slang, homophonic ciphers, and platform-specific idioms) rather than surface token patterns. We conduct a systematic empirical analysis demonstrating that standard metrics including COMET, XCOMET, and BERTScore exhibit near-zero or negative correlation with human cultural judgments, and even display a severity inversion in which scores increase as translation quality deteriorates. We further show that this failure extends to large language model judges: Qwen3-235B achieves Cohen's kappa of only 0.162, revealing that the bottleneck is not reasoning capacity but cultural grounding: models lack the domain-specific cultural knowledge needed to identify which aspects of a translation require scrutiny. To address this, we propose CuRIL, a reinforcement learning framework that internalizes cultural reasoning: cultural annotations are prepended inside the model's reasoning, excluded from policy gradients via a token-level loss mask, and injected with a probability that decays to zero over training, progressively forcing autonomous cultural judgment. On a 1,444-sample human-annotated social media translation benchmark, Qwen3-8B trained with CuRIL achieves Cohen's kappa 0.370 and Exact Match accuracy of 45.22%, approaching Gemini-3.1-Pro with 30x fewer parameters and surpassing models up to 235B in scale. We further demonstrate that our judge produces reliable reward signals for downstream translation optimization, reducing the low-quality translation rate by over 20 percentage points under independent human evaluation.

发表机构

  • Zhejiang University(浙江大学)
  • Xiaohongshu Inc.(小红书科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

↑