rSIM: Incentivizing Reasoning Capabilities of LLMs via Reinforced Strategy Injection
rSIM: 通过强化策略注入激励大语言模型的推理能力
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) ; University of Toronto(多伦多大学) ; University of Alberta(阿尔伯塔大学)
专题命中 规划推理 :reasoning(title,abstract);CoT(abstract);planning(abstract);分类 cs.AI
AI总结 rSIM通过强化策略注入机制,使LLM具备推理能力,并在实验中显著提升模型性能。
Comments 14 pages, 6 figures. Accepted to the ACL ARR July