发表机构
Xiaohongshu Inc.; Tianjin University(小红书公司; 天津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出统一四系数逐Token门控参数化,统一EOPD与ToDi,在TweetEval上优于单通道限制,作为OPD门控比较的共享坐标系。
AI 中文摘要
逐Token的前向/反向KL损失门控已成为在线策略知识蒸馏(OPD)的标准技术,但现有方法如EOPD(Jin等人,2026)和ToDi(Jung等人,2025)各自固定单一门控信号和单一门控方向,且两者从未被直接比较过。我们引入了一个四系数参数化形式 λ_t = σ(a * h_t + b * u(x) + c + d * gap_t),其中EOPD和ToDi的方向对齐代理表现为一维(1D)限制,并增加了多通道组合和显式偏置作为额外自由度。在TweetEval(Barbieri等人,2020)的情感与仇恨任务上,使用Qwen3-32B教师模型和Qwen3-4B学生模型,完整家族中的配置在36个可比单元中的33个中达到了比匹配幅度的单通道(仅熵/仅差距)1D限制更高的准确率,并且一个26单元均值匹配隔离实验将动态门控置于有效KL匹配的静态基线之前,在26个单元中的19个中表现更优。由于单元共享训练数据、模型和参数子结构,我们将这两项计数报告为探索性聚合方向证据,而非独立假设检验。针对该扫描选出的九个头条比较(包括第三个任务——冒犯性语言)进行了定向三种子配对复制,结果方向一致,但个体上小于单种子估计,且在n=3时不显著。因此,我们主要将该参数化作为比较短输出分类OPD中逐Token门控设计的共享坐标系来呈现。
英文摘要
Per-token gating of forward/reverse KL losses has become a standard technique for on-policy knowledge distillation (OPD), but existing methods such as EOPD (Jin et al., 2026) and ToDi (Jung et al., 2025) each fix a single gating signal and a single gating direction, and the two have never been compared directly. We introduce a four-coefficient parameterization lambda_t = sigma(a * h_t + b * u(x) + c + d * gap_t) in which direction-aligned proxies of EOPD and ToDi appear as one-dimensional (1D) restrictions, and which adds multi-channel composition and an explicit bias as further degrees of freedom. On TweetEval (Barbieri et al., 2020) emotion and hate, with a Qwen3-32B teacher and a Qwen3-4B student, configurations in the full family reach higher accuracy than the matched-magnitude single-channel (entropy-only / gap-only) 1D restrictions in 33 of 36 comparable cells, and a 26-cell mean-match isolation experiment places dynamic gating ahead of effective-KL-matched static baselines in 19 of 26 cells. Because cells share training data, models, and parameter substructure, we report both counts as exploratory aggregate directional evidence rather than as independent hypothesis tests. Targeted three-seed paired replications of the nine headline comparisons singled out by that sweep -- including a third task, offensive -- are directionally consistent, but individually smaller than the single-seed estimates and not significant at n=3. We therefore present the parameterization primarily as a shared coordinate system for comparing per-token gating designs in short-output classification OPD.
CommentsAccepted at the Findings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026 Findings)