arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Wnuan:针对专有企业知识的问答模型分阶段后训练方法

Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge

Xiaofeng Shi, Xiaosong Qiu, Wenxin Ma, Qian Kou, Yiming Pan, Longbin Yu, Ying Liu, Haiping Wang, Hua Zhou

arXiv 2608.01862首次发表:更新:

发表机构

Beijing Academy of Artificial Intelligence (BAAI)(北京人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Wnuan是针对专有企业知识的问答模型分阶段后训练方法,通过三阶段流程提升WnuanBench的可接受答案率,同时需控制通用能力的损耗。

AI 中文摘要

企业问答要求模型在获取专有知识的同时不丢失通用能力。我们提出Wnuan,一个三阶段流程:从文档构建面向任务的监督信号,结合通用数据重放进行监督微调(SFT),针对残差误差应用强化学习。在含707个问题的WnuanBench数据集上,主32B路径的可接受答案率(AAR)从适配前的52.76%提升至SFT后的80.06%、强化学习后的91.51%。在匹配的100次更新协议下,残差误差采样的表现比全池采样、规模匹配的随机采样分别高出3.11和2.97个百分点;两类对比的源聚类自举区间均保持为正,同域验证集也维持该排序。该路径下通用基准平均得分下降5.17个百分点,降幅集中在指令跟随任务。自动评估集成与权威领域专家对分层Wnuan-Inst响应样本的一致性达90.5%,这些结果明确了企业分阶段适配的收益及通用能力损耗。

英文摘要

Enterprise question answering requires models to acquire proprietary knowledge without discarding general capabilities. We present Wnuan, a three-stage pipeline that constructs task-oriented supervision from documents, performs supervised fine-tuning with general-data replay, and applies reinforcement learning to residual errors. On the 707-question WnuanBench, the primary 32B route raises acceptable-answer rate (AAR) from 52.76% before adaptation to 80.06% after SFT and 91.51% after RL. Under a matched 100-update protocol, residual-error sampling outperforms full-pool and size-matched random sampling by 3.11 and 2.97 points, respectively. Source-cluster bootstrap intervals remain above zero for both contrasts, and a same-domain validation set preserves the ordering. The general-benchmark average decreases by 5.17 points across the route, concentrated in instruction following. The automatic evaluation ensemble agrees with an authoritative domain expert on 90.5% of a stratified Wnuan-Inst response sample. These results characterize both the gains and the general-capability cost of staged enterprise adaptation.

Comments22 pages, 10 figures. Includes Supplementary Appendices A--L. Xiaofeng Shi and Xiaosong Qiu contributed equally

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑