发表机构
Beijing Academy of Artificial Intelligence (BAAI)(北京人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Wnuan是针对专有企业知识的问答模型分阶段后训练方法,通过三阶段流程提升WnuanBench的可接受答案率,同时需控制通用能力的损耗。
AI 中文摘要
企业问答要求模型在获取专有知识的同时不丢失通用能力。我们提出Wnuan,一个三阶段流程:从文档构建面向任务的监督信号,结合通用数据重放进行监督微调(SFT),针对残差误差应用强化学习。在含707个问题的WnuanBench数据集上,主32B路径的可接受答案率(AAR)从适配前的52.76%提升至SFT后的80.06%、强化学习后的91.51%。在匹配的100次更新协议下,残差误差采样的表现比全池采样、规模匹配的随机采样分别高出3.11和2.97个百分点;两类对比的源聚类自举区间均保持为正,同域验证集也维持该排序。该路径下通用基准平均得分下降5.17个百分点,降幅集中在指令跟随任务。自动评估集成与权威领域专家对分层Wnuan-Inst响应样本的一致性达90.5%,这些结果明确了企业分阶段适配的收益及通用能力损耗。
英文摘要
Enterprise question answering requires models to acquire proprietary knowledge without discarding general capabilities. We present Wnuan, a three-stage pipeline that constructs task-oriented supervision from documents, performs supervised fine-tuning with general-data replay, and applies reinforcement learning to residual errors. On the 707-question WnuanBench, the primary 32B route raises acceptable-answer rate (AAR) from 52.76% before adaptation to 80.06% after SFT and 91.51% after RL. Under a matched 100-update protocol, residual-error sampling outperforms full-pool and size-matched random sampling by 3.11 and 2.97 points, respectively. Source-cluster bootstrap intervals remain above zero for both contrasts, and a same-domain validation set preserves the ordering. The general-benchmark average decreases by 5.17 points across the route, concentrated in instruction following. The automatic evaluation ensemble agrees with an authoritative domain expert on 90.5% of a stratified Wnuan-Inst response sample. These results characterize both the gains and the general-capability cost of staged enterprise adaptation.
Comments22 pages, 10 figures. Includes Supplementary Appendices A--L. Xiaofeng Shi and Xiaosong Qiu contributed equally