arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自信但错误:低资源自动译后编辑的约束解码诊断

Confident but Wrong: A Constrained Decoding Diagnostic for Low-Resource Automatic Post-Editing

Isuru Wijesiri, Nisansa de Silva, Kavindu Warnakulasuriya, Aloka Fernando, Surangika Ranathunga

arXiv 2609.29680首次发表:更新:

发表机构

WSO2; University of Moratuwa; National University of Singapore; Massey University(WSO2; 莫拉图瓦大学; 新加坡国立大学; 梅西大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对低资源自动译后编辑,提出一种黑盒推理时诊断方法,通过编辑距离惩罚曲线和置信度约束排序区分训练不足与数据不一致,揭示两种失败模式,并指导改进。

AI 中文摘要

低资源语言(LRLs)的自动译后编辑(APE)常常无法改进机器翻译(MT),而仅凭分数无法说明原因:是更多训练会有帮助,还是训练数据过于不一致而无法学习。我们提出了一种黑盒、推理时的诊断方法,无需重新训练或标注即可区分这两种情况。该方法通过变化编辑距离惩罚参数 $\lambda$,驱动模型从自由编辑转向复制MT,并读取两个信号:(1)翻译编辑率(TER)随 $\lambda$ 变化的曲线形状——若模型编辑能降低错误则呈U形,若无编辑能降低错误则单调递减;(2)约束变体的排序,这些变体在不同程度上信任模型置信度,以显示置信度是否跟踪编辑质量。在英语-僧伽罗语上的仅解码器模型和编码器-解码器模型中,该诊断揭示了两种与异质后编辑信号作为根本原因相一致的失败模式:二元崩溃(模型复制MT或进行脱靶编辑)和自信误校准(我们测试的置信度信号无法区分有用编辑与不必要编辑)。该模式在英语-马拉地语和英语-泰米尔语上同样成立,失败模式与后编辑分布相关,而非MT质量或语言族。除诊断外,曲线形状为从业者提供了具体的下一步行动;在有利情况下,静态约束可实现免费的推理时准确性提升。我们发布了首个英语-僧伽罗语(约66k)和新的英语-泰米尔语(约39k)APE数据集,并附带所有代码。

英文摘要

Automatic Post-Editing (APE) for low-resource languages (LRLs) often fails to improve Machine Translation (MT), and the score alone cannot say why: whether more training would help, or whether the training data is too inconsistent to learn from. We introduce a black-box, inference-time diagnostic that tells these two cases apart without retraining or annotation. It varies an edit-distance penalty $λ$ that drives the model from free editing towards copying the MT, and reads two signals: (1) the shape of the Translation Edit Rate (TER)-vs-$λ$ curve, U-shaped if edits from the model reduce error and monotonically decreasing if none does; and (2) the ordering of constraint variants that trust model confidence to increasing degrees, which shows whether confidence tracks edit quality. Across decoder-only and encoder-decoder models on English-Sinhala, the diagnostic exposes two failure modes consistent with a heterogeneous post-edit signal as the underlying cause: Binary Collapse, where the model copies the MT or makes off-target edits, and Confident Miscalibration, where the confidence signals we test do not separate useful edits from unnecessary ones. The pattern holds on English-Marathi and English-Tamil, with the failure modes tracking the post-edit distribution rather than MT quality or language family. Beyond diagnosis, the curve shape prescribes a concrete next step for practitioners; in the favorable case, a static constraint yields a free inference-time accuracy gain. We release the first English-Sinhala (~66k) and a new English-Tamil (~39k) APE datasets with all code.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑