arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于低资源印度东北语言的BM25增强多示例翻译

BM25-Augmented Many-Shot Translation for Low-Resource North-Eastern Indian Languages

Aashish Dhawan, Christopher Driggers-Ellis, Dzmitry Kasinets, Christan Grant, Daisy Zhe Wang

arXiv 2608.13722首次发表:更新:

发表机构

University of Florida(佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对WMT26低资源印度东北语言翻译任务,采用BM25检索示例、Gemini 2.5 Flash翻译的无微调方案,通过网格搜索确定最优参数,实现英语与11种印度东北语言的双向翻译。

AI 中文摘要

本文介绍了佛罗里达大学Gators团队提交给WMT26低资源印度语言翻译共享任务的方案。我们将AmericasNLP 2026系统中的检索增强多示例翻译流水线进行适配,用于英语与11种印度东北语言之间的双向翻译。推理阶段,BM25从特定语言的训练库中检索最相似的平行示例,Gemini 2.5 Flash在这些示例的条件下翻译输入,无需进行模型微调。训练库结合了官方WMT26数据与Samanantar等公开语料库及过往WMT共享任务发布的语料。我们在全部22种语言-方向对上对检索数量r和开发示例数量d进行网格搜索,为每个提交选择最优配置。

英文摘要

This paper describes the University of Florida Gators submission to the WMT26 Low-Resource Indic Language Translation shared task. We adapt the retrieval-augmented many-shot translation pipeline from our AmericasNLP 2026 system to translate between English and eleven North-Eastern Indian languages in both directions. At inference time, BM25 retrieves the most similar parallel examples from a language-specific training bank, and Gemini 2.5 Flash translates the input conditioned on these examples. No model fine-tuning is involved. Training banks combine official WMT26 data with publicly available corpora such as Samanantar and prior WMT shared task releases. A grid search over retrieval count r and development exemplar count d across all 22 language-direction pairs selects the best configuration for each submission.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑