TUDUM:面向Qwen3.5-27B的土耳其语思维推理流水线
TUDUM: A Turkish-Thinking Reasoning Pipeline for Qwen3.5-27B
- Department of Artificial Intelligence and Data Engineering(人工智能与数据工程系)
- Faculty of Engineering(工程学院)
- Ankara University(安卡拉大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出TUDUM流水线,通过SFT和GRPO强化学习将Qwen3.5-27B适配为土耳其语思维模型,使推理过程本身使用土耳其语,而非仅输出本地化答案。
AI中文摘要:
本文提出TUDUM(Türkçe Düşünen Üretken Model),一个将Qwen系列27B思维模型适配为土耳其语推理的项目流水线。核心问题不仅是让模型用土耳其语回答土耳其语提示,还要使显式推理轨迹本身使用土耳其语。思维模型可能会将土耳其语提示翻译成以英语为中心的内部或可见草稿,主要用英语解决问题,仅本地化最终答案。TUDUM则将生成的块视为可训练行为。该流水线从项目基础检查点unsloth/Qwen3.5-27B开始,使用LoRA适配器对15,991个土耳其语推理示例进行监督微调(SFT),然后在代理过滤的土耳其语数学环境上应用GRPO系列强化学习。结果喜忧参半。SFT使模型更简洁,推理行为更一致地使用土耳其语,平均响应长度和思维耗尽大幅减少,但降低了基准准确率。RL恢复了一些数学性能,特别是在最佳早期检查点上的AIME24,但并未统一提升所有基准,也未在报告的Macro-6平均值上超过基础模型。因此,贡献最好被定义为技术上诚实的土耳其语思维推理流水线和评估,而非声称达到最先进的土耳其语推理水平。发布的step-50模型已公开可用。
英文摘要:
This paper presents TUDUM (Türkçe Düşünen Üretken Model), a project pipeline for adapting a Qwen-family 27B thinking model toward Turkish reasoning. The central problem is not only to answer Turkish prompts in Turkish, but to make the explicit reasoning trace itself Turkish. A thinking model may translate a Turkish prompt into an English-centered internal or visible scratchpad, solve the problem mostly in English, and only localize the final answer. TUDUM instead treats the generated <think>...</think> block as a trainable behavior. The pipeline starts from the project base checkpoint unsloth/Qwen3.5-27B, applies supervised fine-tuning (SFT) on 15,991 Turkish reasoning examples using LoRA adapters, and then applies GRPO-family reinforcement learning on a proxy-filtered Turkish mathematics environment. The results are mixed. SFT made the model shorter and more consistently Turkish in its reasoning behavior, with large reductions in average response length and thinking exhaustion, but reduced benchmark accuracy. RL recovered some mathematical performance, especially AIME24 at the best early checkpoint, yet did not uniformly improve all benchmarks and did not exceed the base model on the reported Macro-6 average. The contribution is therefore best framed as a technically honest Turkish-thinking reasoning pipeline and evaluation, not as a claim of state-of-the-art Turkish reasoning. The released step-50 model is publicly available.