发表机构
Sakarya University; TU Clausthal; University of Luxembourg; Luxembourg Institute of Science and Technology(萨卡里亚大学; 克劳斯塔尔工业大学; 卢森堡大学; 卢森堡科学技术研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究将翻译视为多智能体探索的决策空间,针对土耳其语-叙利亚阿拉伯语翻译开展实证研究,提出可解释低资源方言生成的框架,轻量稳定化可提升方言标记使用量并降低结构不稳定性。
AI 中文摘要
神经机器翻译(NMT)系统通常为每个输入生成单一输出,掩盖了多语言解码过程中隐含的替代决策轨迹。这种不透明性在低资源方言场景中尤为突出,因为多种语言上有效的实现可能在词汇真实性、语域和结构稳定性方面存在差异。我们提出将翻译重新定义为自主翻译智能体探索的结构化决策空间,不再仅分析单一输出,而是将不同的翻译路径建模为在共享多语言主干上运行的智能体,智能体间的分歧被视为可解释的行为信号而非错误。我们针对土耳其语-叙利亚阿拉伯语翻译开展实证研究,使用三类智能体:(1)零样本直接翻译;(2)通过轻量微调实现的方言稳定翻译;(3)通过英语的枢轴翻译。评估在5000条对话语句上进行,稳定化训练则使用来自电视对话和MADAR-Turk资源的额外5000对土耳其语-叙利亚阿拉伯语句。我们未针对常规性能指标优化,而是通过方言标记频率、与标准阿拉伯语的词汇邻近度及结构方差量化结构化行为位移:轻量稳定化使方言标记使用量近乎翻倍,从0.2266增至0.4988,同时显著降低结构不稳定性;枢轴调解引入归一化压力和可测量的压缩效应;零样本翻译则表现出最高的决策方差。我们认为,智能体间的翻译分歧揭示了多语言模型中潜在的决策灵活性,并为低资源方言生成提供了一个原则性的可解释性框架。
英文摘要
Neural machine translation (NMT) systems typically produce a single output per input, obscuring the alternative decision trajectories implicitly available within multilingual decoding. This opacity becomes particularly problematic in low-resource dialect settings, where multiple linguistically valid realizations may differ in lexical authenticity, register, and structural stability. We propose reframing translation as a structured decision space explored by autonomous translation agents. Instead of analyzing a single output, we model distinct translation pathways as agents operating over a shared multilingual backbone. Inter-agent divergence is treated not as error but as an interpretable behavioral signal. We conduct an empirical study on Turkish--Syrian Arabic translation using three agents: (1) zero-shot direct translation, (2) dialect-stabilized translation via lightweight fine-tuning, and (3) pivot translation through English. Evaluation is performed on 5,000 dialogue sentences, while stabilization is trained on 5,000 additional Turkish--Syrian sentence pairs drawn from television dialogue and MADAR-Turk resources. Rather than optimizing for conventional performance metrics, we quantify structured behavioral displacement using dialect marker frequency, lexical proximity to standardized Arabic, and structural variance. Lightweight stabilization nearly doubles dialect marker usage, increasing it from 0.2266 to 0.4988, while significantly reducing structural instability. Pivot mediation introduces normalization pressure and measurable compression effects, whereas zero-shot translation exhibits the highest decision variance. We argue that translation divergence across agents reveals latent decision flexibility within multilingual models and we provide a principled interpretability framework for low-resource dialect generation.