Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning
专题命中 测试时计算 :reasoning(abstract);分类 cs.AI、cs.LG
Comments preprint
AI 大模型
大模型数学、逻辑、规划、多步推理和测试时计算能力。
专题命中 测试时计算 :reasoning(abstract);分类 cs.AI、cs.LG
Comments preprint