发表机构
BRAC University(BRAC大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对医疗 API 路由任务发现小模型知识蒸馏存在双峰种子崩溃等失败模式,单种子评估无法检测,仅 progressive_kd 和 rank_kd 可避免崩溃。
AI 中文摘要
功能路由是指根据自然语言请求从固定目录中选择正确 API 调用的部署问题,在该场景中,小型学生模型颇具吸引力,但知识蒸馏(KD)的增益通常基于单种子报告,而在该规模下种子方差是未知的。在一项包含 740 个实例的医疗 API 路由任务中,使用 15 亿参数的 Qwen 学生模型和 200 亿参数的教师模型,我们针对有监督交叉熵对比了 8 种 KD 变体,对关键配置使用 3 至 6 个种子。我们发现:(i)每种子标准偏差范围为 2.8 至 48.7 个百分点,吞噬了所有声称的低于 5 个百分点的 KD 增益;(ii)7 种 KD 变体中有 3 种表现出双峰崩溃,至少 1/3 至 5 个种子的准确率低于 55%,其余种子正常训练,还有 1 种变体表现出升高的方差;(iii)崩溃具有 distinct 模式——ce_kd 和 ce_paraphrase 模式为错误函数选择,reasoning_kd 模式为此前未记录的输出截断模式,模型会生成推理内容但在输出函数名称前终止,准确率为 0.9%;(iv)仅 progressive_kd 和 rank_kd 在观察到的种子中避免了崩溃,标准偏差 σ ≤ 3.9 个百分点;(v)输入丰富带来的 naive 跨分割 +3.78 个百分点增益,在受控的分割内多种子重测下反转至 -2.70 个百分点。因此,单种子评估无法检测小模型 KD 中的核心失败模式。
英文摘要
Function routing -- selecting the correct API call from a fixed catalog given a natural-language request -- is a deployment problem where small students are attractive but knowledge distillation gains are typically reported single-seed, at scales where seed variance is unknown. On a 740-instance healthcare API routing task with a 1.5B Qwen student and a 20B teacher, we compare eight KD variants against supervised cross-entropy, using three to six seeds for key configurations. We find: (i) per-seed standard deviation ranges from 2.8 to 48.7 percentage points, swallowing every claimed KD gain below five points; (ii) three of seven KD variants exhibit bimodal collapse, with at least one in three to five seeds falling below 55% accuracy while the others train normally, and a fourth showing elevated variance; (iii) collapse has distinct modes -- wrong-function selection for ce_kd and ce_paraphrase, and a previously undocumented output-truncation mode for reasoning_kd, where the model emits reasoning but terminates before producing a function name (0.9% accuracy); (iv) only progressive_kd and rank_kd avoid collapse across observed seeds, with sigma <= 3.9 pp; (v) a naive cross-split +3.78 pp gain from input enrichment reverses to -2.70 pp under controlled within-split multi-seed re-testing. Single-seed evaluation is therefore unable to detect central failure modes in small-model KD.