arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

跨不同数据集对英语-泰米尔语和泰米尔语-英语的大语言模型和基于Transformer的机器翻译的系统分析

Systematic Analysis of Large Language Models and Transformer-Based Machine Translation for English-Tamil and Tamil-English Across Diverse Datasets

Sriharshaa S, Sangeetha Sivanesan, Jaya Nirmala S

arXiv 2607.24515首次发表:更新:

AI 中文总结

研究针对泰米尔语等低资源语言机器翻译挑战,在多数据集上评估多种模型,用BLEU等指标分析其性能,通过注意力分析增强可解释性,证明上下文提示可实现少样本翻译并与监督方法比较,揭示数据集及领域匹配对模型性能的影响。

AI 中文摘要

泰米尔语等低资源语言的机器翻译面临挑战,主要源于平行数据有限、领域差异大及形态复杂。本研究对多个数据集(NTREX、EnTamV2、WikiMatrix和PMIndia)上的多语言翻译模型在英-泰及泰-英翻译的性能进行全面评估。使用BLEU和chrF指标评估监督式NMT系统、NLLB和mBART,分析其在不同质量和领域数据上的表现。通过可视化源文本与译文标记对齐进行注意力分析以增强模型可解释性。还证明使用上下文提示能借助TamilLaMA模型进行少样本翻译,并与监督方法定性比较。结果表明数据集质量及其与领域的匹配度极大影响模型性能,基于注意力的机制有助于解释,少样本大语言模型仍能生成结构连贯的泰米尔语翻译。

英文摘要

The challenge of Machine Translation for low resource languages such as Tamil is primarily caused by the restricted amount of parallel data for these languages, as well as their substantial amount of domain variation and morphological complexity. This research presents the comprehensive evaluation of the performance of several multilingual translation models on English-Tamil and Tamil-English translations across multiple datasets: NTREX, EnTamV2, WikiMatrix and PMIndia. This study evaluates supervised NMT systems, NLLB and mBART, using both the BLEU and chrF metric, and examines how these systems perform on data of different quality levels and domains. This performs an attention-based analysis to increase model interpretability by visualising the alignments of tokens in an English source text and their Tamil translations and vice-versa to provide insight into how they make translations. This study also demonstrates that using in-context prompting can provide an excellent way to perform a few-shot translation of English to Tamil and Tamil-English using a Tamil capable TamilLaMA model, and compare this to supervised approaches qualitatively. These findings show that the quality of the datasets and their alignment with the domain will greatly affect the performance of the model, that attention-based mechanisms can aid in explain ability, and that few-shot large language models can still produce structurally coherent translations of Tamil.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑