arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

事实胜于虚构:僧伽罗语-英语神经机器翻译中病理性幻觉的检测

Fact over Fiction: Detection of Pathological Hallucinations in Sinhala-to-English Neural Machine Translation

Navam Obeysekara, Nevidu Jayatilleke

arXiv 2610.11389首次发表:更新:

发表机构

School of Computing, Informatics Institute of Technology; University of Moratuwa(信息技术学院计算机学院; 莫拉图瓦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对僧伽罗语-英语NMT的幻觉问题,提出无参考检测框架,构建4.5万样本合成数据集,微调mDeBERTa-v3实现标记级F1达0.841,经基准测试验证了检测器的有效性。

AI 中文摘要

神经机器翻译(NMT)模型虽能生成流畅度极高的输出,但仍易出现幻觉,即译文自然流畅却与源文本语义无关。在僧伽罗语-英语这类低资源场景中,跨语言对齐能力薄弱会加剧这种幻觉问题。本文针对该语言对提出一种无参考的幻觉检测框架。我们通过包含5种语言动机的损坏策略的概率链生成了45000个样本的合成数据集,该数据集带有语义救援机制,利用字符级相似度区分幻觉与形态变体。我们对mDeBERTa-v3进行微调以实现标记级序列标注,在源不相交测试集上,3个随机种子的标记级F1值达0.841±0.001;同时研究了整合神经风险评分、序列对数概率和跨语言语义嵌入(LaBSE)的三信号集成方法。源消融控制实验显示,检测器依赖僧伽罗语源文本而非损坏过程的表面伪影:打乱或移除源文本会使句子级AUROC从0.970降至随机水平。我们对涵盖5个模型家族的8个NMT系统进行基准测试,发现检测器触发率在不同架构间相差一个数量级。

英文摘要

Neural Machine Translation (NMT) models, while capable of producing highly fluent outputs, remain vulnerable to hallucinations, which are translations that are natural yet semantically unrelated to the source. This vulnerability is acute in low-resource settings like Sinhala-to-English, where weak cross-lingual alignment leads to hallucinations. This paper introduces a framework for reference-free hallucination detection in this language pair. We present a 45,000-sample synthetic dataset generated through a probabilistic chain of five linguistically motivated corruption strategies, with a semantic rescue mechanism that uses character-level similarity to distinguish hallucinations from morphological variants. We fine-tune mDeBERTa-v3 for token-level sequence labelling, reaching a token-level F1 of 0.841 +/- 0.001 over three seeds on a source-disjoint test set, and study a three-signal ensemble integrating neural risk scores, sequence log-probabilities, and cross-lingual semantic embeddings (LaBSE). A source-ablation control shows that the detector relies on the Sinhala source rather than on surface artefacts of the corruption process: shuffling or removing the source reduces sentence-level AUROC from 0.970 to chance. We benchmark eight NMT systems spanning five model families and find that detector firings vary by an order of magnitude across architectures.

Comments11 pages, 1 figure, 7 tables, Accepted paper at the 13th Conference on Computational Linguistics and Speech Processing (ROCLING) 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑