arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向领域后训练的自专用教师

Self-Specialized Teachers for Domain Post-Training

Yifei Li, Rongman Xu, Lingling Zhang, Muye Huang, Zihan Ma, Jiashuai Liu, Hang Yan, Heng Wang

arXiv 2608.28647首次发表:更新:

AI 中文总结

针对无通用重放语料的领域后训练问题,提出SSTD方法,在保留目标域性能提升的同时,提升通用套件分数4.8-5.0点,且适配不同模型规模与主干。

AI 中文摘要

仅目标域的后训练可提升特定领域的性能,但会降低通用基础模型在适配前习得的行为。本文研究当目标域数据可用但无代表性重放语料时的该问题,提出自专用教师蒸馏(SSTD)这一两阶段流程:首先将基础模型的副本训练为领域教师,再基于学生自身采样的前缀,将教师的词元分布蒸馏给学生。教师训练结合标准目标监督、感知基础模型的关键词元加权,以及对冻结基础模型的分布对齐;在线策略蒸馏则将领域反馈置于学生推理时可能遇到的状态。在金融数值推理、医学问答、法律条款识别任务上,SSTD保留了直接微调的大部分目标域提升,同时在报告的操作点上,使评估通用套件的平均分数提升4.8至5.0个百分点。该模式在Qwen3不同规模及Gemma主干上均成立,且SSTD无需外部教师或通用重放数据。

英文摘要

Target-only post-training can improve performance in a specialized domain while degrading behaviors that a general-purpose base model acquired before adaptation. We study this problem when target-domain data are available but a representative replay corpus is not. We propose self-specialized teacher distillation (SSTD), a two-stage procedure that first trains a copy of the base model into a domain teacher, then distills its token distribution to a student on prefixes sampled from the student itself. Teacher training combines standard target supervision with base-aware key-token weighting and distribution alignment to the frozen base model; on-policy distillation then places domain feedback on states the student can encounter at inference time. On financial numerical reasoning, medical question answering, and legal holding identification, SSTD retains much of the target improvement of direct fine-tuning while improving the mean score on the evaluated general suite by 4.8--5.0 points at the reported operating point. The pattern persists across Qwen3 sizes and on Gemma backbones. SSTD requires neither an external teacher nor general replay data.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑