arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GSM-Plus-BN:大语言模型中孟加拉语数学推理的基于扰动的基准测试

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models

Bidyarthi Paul, Nahida Jannat Mayouree, Md. Asif Karim, Sagar Chandra Nath, Swastika Kundu

arXiv 2607.13248首次发表:更新:

发表机构

Southeast University; Ahsanullah University of Science and Technology(东南大学; 阿山努拉科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对大语言模型中孟加拉语数学推理评估缺乏的问题,引入GSM-Plus-BN数据集,用9000个样本评估六个开源LLMs,对比标准提示和思维链提示,发现大模型鲁棒性好,思维链提示有提升,但与英语基准有差距,为相关研究提供新资源和基线。

AI 中文摘要

大语言模型(LLMs)中数学推理评估主要集中在英语等资源丰富的语言上。这给像孟加拉国这样有超过2.3亿人说孟加拉语的语言多样化地区的人工智能公平发展和部署造成重大障碍。此前孟加拉语数学推理相关工作极少,且无系统基准测试。本研究引入GSM-Plus-BN,一个源自英语GSM-Plus基准并经人工翻译验证的新型扰动孟加拉语数学数据集。使用9000个评估样本(含1000个种子问题和8000个扰动变体)对六个开源LLMs进行评估。实验结果表明,在标准提示下GPT-OSS-20B种子问题准确率最高达96.08%,大模型在扰动类型上表现出更好的鲁棒性。思维链提示比标准提示显著提升多数模型推理能力,但所有模型与英语基准相比仍有差距。本研究提供GSM-PLUS-BN作为新资源和基线,为未来孟加拉语数学推理研究做出了基础性贡献。

英文摘要

The evaluation of mathematical reasoning in large language models (LLMs) has predominantly focused on high-resource languages like English. This has created a significant barrier to the equitable development and deployment of AI in linguistically diverse regions such as Bangladesh, where over 230 million people speak Bengali. Despite this global significance, there has been minimal prior work on mathematical reasoning in Bengali and no existing research that systematically benchmarks a perturbated Bengali mathematical dataset, leaving a critical void in assessing model robustness and true comprehension beyond pattern recognition. This study addresses this gap by introducing GSM-Plus-BN, a novel perturbated Bengali mathematical dataset derived from the English GSM-Plus benchmark and verified by human translators. We evaluate six open-source LLMs Qwen3-32B, Llama-3.1-8B-Instant, Llama-3.3-70B-Versatile, Llama-4-Scout-17B-16E-Instruct, GPT-OSS-120B, and GPT-OSS-20B using a benchmark of 9,000 evaluation samples comprising 1,000 seed questions and 8,000 perturbed variants under both Standard Prompting and Chain-of-Thought (CoT) Prompting. Experimental results show that GPT-OSS-20B achieves the highest seed question accuracy of 96.08% under Standard Prompting, while larger models such as Llama-3.3-70B and GPT-OSS-120B demonstrate superior robustness across perturbation types. Furthermore, CoT prompting substantially improves reasoning for most models compared to Standard Prompting, yet a notable performance gap persists across all models relative to their English benchmarks, underscoring the inherent difficulty of perturbed Bengali text. This research makes a foundational contribution by providing GSM-PLUS-BN as a new resource and baseline for future Bengali mathematical reasoning research.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑