发表机构
University at Albany(奥尔巴尼大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究构建含GPT-5.4推理依据的孟加拉语数学数据集MathShikkha,微调4个4B-7B模型,发现CoT监督对域内强骨干无增益但提升弱模型,域外及BanglaMATH基准上均优于仅答案,核心益处为语言贴合度、可审核推理及域外鲁棒性。
AI 中文摘要
数学推理在孟加拉语等低资源语言中仍具挑战性。本研究探究教师生成的孟加拉语思维链(CoT)监督是否能超越普通监督微调带来增益。我们构建了\textsc{MathShikkha}数据集,这是一个含GPT-5.4生成推理依据的孟加拉语数学推理数据集,并在匹配协议下微调了4个4B至7B规模的学生模型,该协议中仅答案与CoT条件共享数据划分、仅响应损失掩码、解码及评分,仅训练目标不同。在域内场景,尽管CoT生成的token数量是仅答案的15至52倍,但对于三个较强骨干模型,CoT未较仅答案微调有显著提升(配对自助法95%置信区间包含0,精确McNemar检验p≥0.17),仅对较弱的4B模型提升了18.56个百分点(p<0.0001)。在更大且经污染审核的BanglaMATH基准上,该模式反转:CoT对所有四个模型均较仅答案监督提升20.1至28.1个百分点(所有p<0.0001)。仅答案微调还使三个模型的域外准确率低于基础模型,而CoT对所有四个模型均保留或提升了域外准确率。由两位共同作者标注者开展的人类研究,经外部专家裁决,Cohen’s κ值为0.76至1.00,发现CoT在推理内容标准上未较基础模型有显著提升;其可测量的效果是目标语言 adherence( adherence 指 adherence to the target language,即目标语言贴合度)和产生可检查的推理。总体而言,推理依据监督的价值取决于骨干模型能力与分布偏移:在本研究场景中,其主要益处是孟加拉语贴合度、可审核推理及域外鲁棒性,而非提升域内推理有效性。
英文摘要
Mathematical reasoning remains challenging in low-resource languages such as Bangla. We study whether teacher-generated Bangla Chain-of-Thought (CoT) supervision provides benefits beyond ordinary supervised fine-tuning. We construct \textsc{MathShikkha}, a Bangla mathematical reasoning dataset with GPT-5.4-generated rationales, and fine-tune four 4B--7B student models under a matched protocol in which answer-only and CoT conditions share data splits, response-only loss masking, decoding, and scoring, differing only in the training target. In-domain, CoT provides no significant improvement over answer-only fine-tuning for three stronger backbones (paired bootstrap 95\% CIs include zero; exact McNemar $p \geq 0.17$), despite generating 15--52$\times$ more tokens, but significantly improves the weaker 4B model by 18.56 points ($p < 0.0001$). On the larger, contamination-audited BanglaMATH benchmark, this pattern reverses: CoT significantly outperforms answer-only supervision for all four models by 20.1--28.1 points (all $p < 0.0001$). Answer-only fine-tuning also reduces out-of-domain accuracy below the base model for three models, whereas CoT preserves or improves it for all four. A human study with two co-author annotators, external-expert adjudication, and Cohen's $κ= 0.76$--$1.00$ finds no significant CoT improvement over the base model on reasoning-content criteria; instead, its measurable effect is target-language adherence and producing inspectable reasoning. Overall, rationale supervision's value depends on backbone capability and distribution shift: in this setting, its main benefits are Bangla adherence, auditable reasoning, and out-of-domain robustness rather than improved in-domain reasoning validity.