发表机构
School of Cyber Science and Technology; Beihang University; School of Information; Renmin University of China(网络科学与技术学院; 北京航空航天大学; 信息学院; 中国人民大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对通用大语言模型在科学健身教练中领域知识不足的问题,提出FitOne系列模型,通过三阶段后训练(持续预训练、监督微调、强化学习)提升专业能力,在ACSM-EP和NSCA-CSCS考试中平均提升最高12.73%。
AI 中文摘要
科学健身教练(SFC)通常由人类专业人士提供,导致成本高昂且难以普及。尽管大语言模型(LLMs)的最新进展显示出更包容的健身教练的巨大潜力,但直接在SFC中部署现有的通用LLMs暴露了关键限制。这些模型往往缺乏足够的领域特定知识整合,导致在复杂SFC场景中表现不佳。在本文中,我们介绍FitOne,一系列健身LLMs(8B和32B参数),旨在提高SFC应用的可靠性和领域专业化。基于Qwen3基础模型,FitOne通过三阶段后训练流水线开发,包括持续预训练、监督微调和强化学习,使用来自严格知识工程的大规模高质量数据集。我们在专业健身认证考试(包括ACSM-EP和NSCA-CSCS)以及通用能力(如知识推理和指令遵循)上对FitOne进行全面评估。实验结果表明,在保持强大通用能力的同时,与Qwen3基础模型相比,FitOne-8B/32B在ACSM-EP和NSCA-CSCS考试上分别实现了平均高达10.09%/9.29%和12.73%/7.01%的提升。此外,深入的消融研究证实了每个训练阶段的必要性,突出了该流水线在平衡领域专业知识增强与通用能力保持方面的有效性。我们相信这项研究将推动LLM系统向更可靠的健身智能发展,并将激发未来关于开发领域特定LLMs的研究。
英文摘要
Scientific Fitness Coaching (SFC) is typically delivered by human professionals, making it costly and inaccessible to many. While recent advances in Large Language Models (LLMs) show considerable promise for more inclusive fitness coaching, directly deploying prevailing general-purpose LLMs in SFC reveals critical limitations. These models often lack sufficient domain-specific knowledge integration, leading to weak performance on complex SFC scenarios. In this paper, we introduce FitOne, a series of fitness LLMs (with 8B and 32B parameters) designed to improve reliability and domain specialization for SFC applications. Built upon the Qwen3 foundation models, FitOne is developed through a three-stage post-training pipeline consisting of continual pre-training, supervised fine-tuning, and reinforcement learning, using large-scale, high-quality datasets derived from rigorous knowledge engineering. We conduct comprehensive evaluations of FitOne on professional fitness certification exams, including ACSM-EP and NSCA-CSCS, as well as general capabilities such as knowledge reasoning and instruction following. Experimental results show that, while retaining strong general capabilities, FitOne-8B/32B achieves average improvements of up to 10.09%/9.29% and 12.73%/7.01% on the ACSM-EP and NSCA-CSCS exams, respectively, compared with the Qwen3 base models. Furthermore, in-depth ablation studies confirm the necessity of each training stage, highlighting the pipeline's effectiveness in balancing domain expertise enhancement with general ability retention. We believe this research advances LLM systems toward more reliable fitness intelligence and will inspire future research on developing domain-specific LLMs.