arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11167cs.LGcs.AI

PIVOT:基于困惑度的KD到RL过渡调度的垂直领域少样本蒸馏方法

PIVOT: Perplexity-Informed KD-to-RL Transition Scheduling for Vertical-Domain Few-Shot Distillation

Heng Li, Yong Zhang, Ning Cheng, Zhigen Li, Yun Zhu, Yanmeng Wang, Shaojun Wang, Jing Xiao

首次发表
浏览论文内容

中文总结 AI 辅助

针对垂直领域少样本分类中KD到RL过渡调度固定的问题,提出PIVOT动态框架,按样本困惑度在OPD与GRPO间路由,在Banking77等数据集上性能优于基线且训练更稳定。

中文摘要 AI 辅助

垂直领域少样本分类对于小型语言模型而言仍具挑战性,因为有限的监督信号使其难以获取领域特定的决策知识。在线策略蒸馏(On-Policy Distillation, OPD)可通过监督学生生成的rollout提升教师引导的适配效果,而基于GRPO的强化学习能进一步优化下游预测。然而,现有的KD到RL流水线通常依赖全局固定的过渡调度,忽略了不同样本在奖励驱动优化前可能需要不同量的教师引导知识获取。我们提出PIVOT(Perplexity-Informed Transition Optimization,基于困惑度的过渡优化),这是一种动态过渡框架,根据教师评估的序列困惑度在OPD和GRPO之间路由样本:将低困惑度样本移至GRPO进行奖励驱动优化,同时将高困惑度样本保留在OPD中以持续获取领域知识。在Banking77和HWU64数据集上的实验表明,在相同的预热后学生优化步数下,PIVOT始终优于持续OPD和全局同步OPD→GRPO基线,实现了更强的下游性能和更稳定的训练动态。

英文摘要

Vertical-domain few-shot classification remains challenging for small language models, as limited supervision makes it difficult to acquire domain-specific decision knowledge. On-Policy Distillation (OPD) can improve teacher-guided adaptation by supervising student-generated rollouts, while GRPO-based reinforcement learning can further refine downstream predictions. However, existing KD-to-RL pipelines typically rely on globally fixed transition schedules, ignoring that different samples may require different amounts of teacher-guided acquisition before reward-driven refinement. We propose PIVOT (Perplexity-Informed Transition Optimization), a dynamic transition framework that routes samples between OPD and GRPO according to teacher-evaluated sequence perplexity. PIVOT moves low-perplexity samples to GRPO for reward-driven refinement while keeping high-perplexity samples under OPD for continued domain knowledge acquisition. Experiments on Banking77 and HWU64 show that PIVOT consistently outperforms continued OPD and globally synchronized OPD$\rightarrow$GRPO baselines under the same number of post-warm-up student optimization steps, achieving stronger downstream performance and more stable training dynamics.

发表机构

  • Ping An Technology (Shenzhen) Co., Ltd.(平安科技(深圳)有限公司)
  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑