探索大型语言模型将罗马尼亚语计算问题翻译为英语
Exploring Large Language Models for Translating Romanian Computational Problems into English
- University of Bucharest(布加勒斯特大学)
- Softbinator
- UiPath
- It Just Works Inc.
- QPillars
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文评估多种LLM在结构化提示与人工监督下将罗马尼亚语信息学竞赛题翻译为英语的准确性和稳定性,并扩充OJI双语数据集,证明其可接近人工翻译质量。
AI中文摘要:
近期研究表明,当数学和计算机科学任务从罗马尼亚语翻译成英语时,大型语言模型(LLM)在这些问题上的表现不如其原始罗马尼亚语格式。准确翻译对于从编程竞赛中的自动翻译到创建高质量教育材料,以及最大限度减少人工翻译中的错误或欺诈等应用都至关重要。本研究表明,若给出结构良好的提示,稳健的大型语言模型(LLM)在翻译较不常见语言时能够保持甚至提升其性能。我们的研究结果表明,在适当监督下,LLM可以可靠地用于IOI(国际信息学奥林匹克)风格任务的自动翻译。我们在多个LLM上评估了若干翻译方法,包括OpenRoLLM、Llama 3.1 8B、Llama 3.2 3B和GPT-4o,并通过重复运行评估其翻译准确性和性能稳定性。此外,我们用准确的英语翻译扩充了OJI(罗马尼亚县级信息学奥林匹克)罗马尼亚语数据集,增强了其在未来LLM训练和评估中的效用。通过详细的句法和语义分析,我们确认在人工监督下,LLM可以成为多语言问题求解的可行解决方案。我们还由一位认证专家评估,比较了LLM与人工译者的翻译质量,凸显了LLM在现实场景中的潜力。
英文摘要:
Recent studies have suggested that large language models (LLMs) underperform on mathematical and computer science tasks when these problems are translated from Romanian into English, compared to their original Romanian format. Accurate translation is critical for applications ranging from automatic translations in programming competitions to the creation of high-quality educational materials, as well as minimizing errors or fraud in human translations. This study shows that robust large language models (LLMs) can maintain or even enhance their performance in translating less common languages when given well-structured prompts. Our findings suggest that LLMs, with appropriate supervision, can be reliably used for the automatic translation of IOI (International Olympiad in Informatics)-style tasks. We evaluate several translation methods across multiple LLMs, including OpenRoLLM, Llama 3.1 8B, Llama 3.2 3B and GPT-4o, assessing their translation accuracy and performance stability through repeated runs. Additionally, we augment the OJI (Romanian County-Level Informatics Olympiad) Romanian dataset with accurate English translations, enhancing its utility for future LLM training and evaluation. Through detailed syntactic and semantic analyses, we confirm that with human oversight, LLMs can serve as a viable solution for multilingual problem-solving. We also compare the translation quality of LLMs against human translators, as evaluated by a certified expert, underscoring the potential of LLMs in realworld scenarios.