话语依存:一种翻译难度的连续判定标准
Discourse Dependency: A Continuous Criterion for Translation Difficulty
浏览论文内容
中文总结 AI 辅助
该研究提出话语依存(DDP)作为翻译难度的连续判定指标,验证其有效性后,将其应用于WMT数据集并对比不同上下文注入策略,发现高DDP片段需长距离上下文,可用于优化机器翻译评估。
中文摘要 AI 辅助
近期对更具挑战性的机器翻译基准的呼吁并未明确难度的定义。我们认为,一个有意义且目前无法测量的维度是指涉范围,即一个片段必须回溯到其文档中以解析所含实体和代词的距离。我们将此形式化为话语依存(Discourse Dependency,DDP),这是一种无度量的源侧指标,由命名实体的再次提及和代词共指关系计算得出。经黄金标准共指关系验证,99.2%的片段出现单侧错误,因此高DDP片段被证实需要长距离上下文。将DDP应用于WMT24++和WMT25显示,两者都严重偏向低DDP片段,而领域标签无法区分这些片段。基于DDP,我们在英韩后编辑设置中比较了五种上下文注入策略,改变上下文大小和选择。随着DDP增长,没有策略能跟上人类后编辑的表现。在DDP≥15的片段中,评分者更偏好人类翻译,而自动指标未记录到差异。当前沿系统的 aggregate 分数达到饱和时,DDP将评估从模型得分的好坏转向它们能覆盖的距离。
英文摘要
Recent calls for harder machine translation benchmarks have not clarified what difficulty should mean. We argue that one meaningful and currently unmeasured axis is referential reach, the distance a segment must look back into its document to resolve the entities and pronouns it contains. We formalize this as discourse dependency (DDP), a metric-free, source-side measure computed from named entity re-mentions and pronominal coreference. Validated against gold coreference, DDP errs one-sidedly in 99.2% of segments, so a high-DDP segment is certified to require long-range context. Applying DDP to WMT24++ and WMT25 shows that both are heavily skewed toward low-DDP segments, which domain labels do not distinguish. Building on DDP, we compare five context injection strategies in an English-Korean post-editing setup, varying context size and selection. As DDP grows, no strategy keeps pace with human post-editing. On segments with DDP >= 15 raters prefer human translations, while automatic metrics register no difference. As frontier systems saturate aggregate scores, DDP shifts evaluation from how well models score to how far they can reach.
发表机构
- AI-Bio Convergence Research Inst.(AI-生物融合研究所)
- School of Software(软件学院)
- Soongsil University(崇实大学)
机构由 AI 辅助整理,请以论文原文为准。