A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning
机构 * Aillis Inc.(Aillis公司) ; Department of Health Policy(健康政策部门) ; Public Health, Graduate School of Pharmaceutical Sciences, The University of Tokyo, Tokyo, Japan(公共卫生,药学研究生院,东京大学,东京,日本) ; Rist Inc.(Rist公司)
专题命中 指令微调 :SFT(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
Comments Presented at ICML 2025 Workshop on The second AI for MATH