AI 中文总结
本研究针对信号数学推理任务,以Qwen2.5-3B-Base为对象,对比两种训练范式与三种强化优化算法,其最佳模型准确率较基准提升超两倍。
AI 中文摘要
通过监督思维链微调及可验证奖励的强化学习进行后训练,已大幅提升了大型语言模型(LLMs)的数学推理能力,但其在信号处理问题中的应用仍相对未被充分探索。本报告研究了将Qwen2.5-3B-Base适配至WirelessMATHBench-XL中研究生水平信号数学问题的强化微调策略,WirelessMATHBench-XL是该领域数学推理的综合基准。我们考察两种训练范式:(i)在WirelessMATHBench-XL上采用可验证奖励的直接强化学习(RL);(ii)在蒸馏后的无线领域思维链语料库上进行监督微调(SFT),随后进行相同的领域特定RL阶段。在两种范式下,我们对Group Relative Policy Optimization(GRPO)、Group Sequence Policy Optimization(GSPO)和Geometric-Mean Policy Optimization(GMPO)进行了基准测试。我们旨在探究领域感知思维链SFT是否能作为后续RL的有效初始化,以及GSPO或GMPO在信号推理任务的稳定性或准确性上是否优于GRPO。我们的最佳模型实现了39.12%的总体准确率,较未训练的Base模型(12.37%)提升超两倍。
英文摘要
Post-training with supervised chain-of-thought fine-tuning and reinforcement learning from verifiable rewards has substantially improved the mathematical reasoning capabilities of large language models (LLMs). However, their application to signal processing problems remains relatively under-explored. This report investigates reinforcement fine-tuning strategies for adapting Qwen2.5-3B-Base to graduate-level signal mathematical problems from WirelessMATHBench-XL, a comprehensive benchmark for mathematical reasoning in this domain. We examine two training paradigms: (i) direct reinforcement learning (RL) on WirelessMATHBench-XL with verifiable rewards; and (ii) supervised fine-tuning (SFT) on a distilled wireless-domain chain-of-thought corpus, followed by the same domain-specific RL stage. Across both paradigms, we benchmark Group Relative Policy Optimization (GRPO), Group Sequence Policy Optimization (GSPO), and Geometric-Mean Policy Optimization (GMPO). We aim to assess whether domain-aware CoT SFT serves as an effective initialization for subsequent RL, and whether GSPO or GMPO offer advantages in stability or accuracy over GRPO for signal reasoning tasks. Our best model achieves an overall accuracy of 39.12\%, representing a more than threefold improvement over the untrained Base model (12.37\%).