arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学生引导的教师蒸馏用于高效LLM任务路由:定位对抗Jev风格System-1分类器

Student-Guided Teacher Distillation for Efficient LLM Task Routing: Positioning Against Jev-Style System-1 Classifiers

Haifeng Wu, Srinivasan Manoharan, Jian Wan, Fangbo Tu, Junhua Zhao, Xin Chen

arXiv 2610.02516首次发表:更新:

发表机构

PayPal(PayPal)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出学生引导教师蒸馏流程,用紧凑ModernBERT生成top-k候选并由DeBERTa-v3重排序,实现高效LLM任务路由,达到77.5%教师一致性,并揭示截断软标签的偏差问题。

AI 中文摘要

零样本分类器对于将用户请求路由到专门的LLM任务非常有用,但针对大型候选集对每个请求进行评分代价高昂:零样本NLI分类器必须为每个标签评估一个前提-假设对,因此成本随分类体系规模线性增长。我们研究了一个针对固定60个LLM任务类别分类体系的学生引导教师蒸馏流程:一个紧凑的ModernBERT分类器在一次前向传播中预测完整类别分布并检索一个小的top-k候选集,一个更大的DeBERTa-v3零样本NLI分类器仅对这些候选进行重排序,而非全部60个标签;由此产生的教师标签迭代改进学生,使其为下一轮生成更尖锐的候选。与极端多标签分类中使用的通用嵌入检索或聚类派生短列表不同,我们的候选生成器在目标分类体系上端到端训练,并且是服务生产流量的同一模型,这使其区别于在候选模型之间路由的LLM路由工作,也区别于并发System-1编码器分类器提案(如TypeSafe AI的Jev和开源Laya项目),这些提案的训练方法未记录或基于RL。我们最好的学生检查点在200个示例的评估集上达到77.5%的教师一致性,初步覆盖率测量显示Coverage@16为91-100%,表明top-k集保留了教师决策相关信息的绝大部分。我们进一步表明,截断的top-k教师分数不应被视为用于KL蒸馏的完整60类软目标:将未截断类别归零会破坏软标签蒸馏所依赖的暗知识,引入系统性偏差而非无害的稀疏近似。完整的评估,包括在保留集上多个k的覆盖率、嵌入检索基线和更大的人工审查测试集,仍在进行中。

英文摘要

Zero-shot classifiers are useful for routing user requests to specialized LLM tasks, but scoring every request against a large candidate set is expensive: a zero-shot NLI classifier must evaluate one premise-hypothesis pair per label, so cost scales linearly with taxonomy size. We study a student-guided teacher distillation pipeline for a fixed taxonomy of 60 LLM task categories: a compact ModernBERT classifier predicts the full category distribution in one forward pass and retrieves a small top-k candidate set, and a larger DeBERTa-v3 zero-shot NLI classifier reranks only those candidates rather than all 60 labels; the resulting teacher labels iteratively improve the student, which produces sharper candidates for the next round. Unlike generic embedding retrieval or clustering-derived shortlists used in extreme multi-label classification, our candidate generator is trained end-to-end on the target taxonomy and is the same model serving production traffic, distinguishing it from LLM-routing work that routes between candidate models, and from concurrent System-1 encoder-classifier proposals (e.g. TypeSafe AI's Jev and the open-source Laya project) whose training methodology is undocumented or RL-based. Our best student checkpoint reaches 77.5% teacher agreement on a 200-example evaluation set, and preliminary coverage measurements show Coverage@16 of 91-100%, suggesting top-k sets retain most of the teacher's decision-relevant information. We further show truncated top-k teacher scores should not be treated as full 60-class soft targets for KL distillation: zeroing untruncated classes destroys the dark knowledge soft-label distillation depends on, introducing systematic bias rather than a harmless sparse approximation. A complete evaluation, including coverage at multiple k on a held-out set, an embedding-retrieval baseline, and a larger human-reviewed test set, remains in progress.

Comments13 pages, 2 figures, 1 table

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑