arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ALRA:用于自回归语言模型基于Logit的预训练蒸馏的自适应局部关系对齐

ALRA: Adaptive Local Relational Alignment for Logit-Based Pre-training Distillation of Autoregressive Language Models

Quang Hoang Trung, Quang Huu Hieu, Nguyen Van Hoang Phuc, Vo Nguyen Le Duy

arXiv 2609.03355首次发表:更新:

发表机构

VJ Technologies; AJ Technologies; Vietnam National University; University of Information Technology(VJ科技; AJ科技; 越南国家大学; 信息技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对自回归语言模型基于Logit的预训练蒸馏忽略令牌相对偏好的问题,提出ALRA框架,结合学生提议与教师指导,在The Pile数据集的9个零样本基准上提升了蒸馏模型的准确率。

AI 中文摘要

自回归语言模型基于Logit的知识蒸馏通常会在整个词汇表上对齐教师模型与学生模型的下一个令牌分布。然而,这种全局目标忽略了可能令牌备选方案之间的相对偏好。现有局部方法通常仅从教师模型或学生模型中选择候选令牌。仅从教师模型选择可能会遗漏学生模型认为可能的令牌,而仅从学生模型选择可能会在训练早期依赖不准确的排名。我们提出自适应局部关系对齐(Adaptive Local Relational Alignment,ALRA),这是一种结合学生提议与教师指导的位置特定框架。在每个有效预测位置,学生模型会提出可能的令牌,同时教师模型的最可能令牌会被包含作为锚点。ALRA会根据教师模型在该候选集内相对于当前批次的概率分布广度,调整所选令牌的数量。自适应局部散度保留质量匹配项,并分别匹配所选令牌区域与剩余词汇区域内的相对令牌分布。与精确的全词汇分解不同,它用单位系数替换了两个条件项的教师质量系数,防止任一条件项仅因其区域的教师概率低而被降权。学生加权成对关系对齐强调具有小学生概率差距的高概率令牌对,并对不可能或明显分离的对赋予较低权重。在The Pile数据集上,使用随机初始化的2亿和5亿参数学生模型,在9个零样本基准测试中,平均准确率分别达到36.62%和37.40%。ALRA比最强的竞争蒸馏基线分别高出0.94和0.83个百分点,且分别比未进行蒸馏的预训练高出2.31和2.91个百分点。

英文摘要

Logit-based knowledge distillation for autoregressive language models usually aligns teacher and student next-token distributions over the entire vocabulary. However, this global objective overlooks relative preferences among likely token alternatives. Existing local approaches often select candidate tokens from either the teacher or the student alone. Teacher-only selection can miss tokens that the student considers likely, while student-only selection can rely on an inaccurate ranking early in training. We propose Adaptive Local Relational Alignment (ALRA), a position-specific framework combining student proposals with teacher guidance. At each valid prediction position, the student proposes likely tokens, while the teacher's most probable token is included as an anchor. ALRA adjusts the number of selected tokens according to how broadly the teacher distributes probability within this candidate set relative to the current batch. Adaptive Local Divergence retains the mass-matching term and separately matches the relative token distributions within the selected and remaining vocabulary regions. Unlike the exact full-vocabulary decomposition, it replaces the teacher-mass coefficients of the two conditional terms with unit coefficients, preventing either term from being downweighted solely because its region has low teacher probability. Student-Weighted Pairwise Relational Alignment emphasizes high-probability token pairs with small student probability gaps and gives less weight to unlikely or clearly separated pairs. Experiments on The Pile with randomly initialized 200M- and 500M-parameter students across nine zero-shot benchmarks yield average accuracies of 36.62% and 37.40%. ALRA exceeds the strongest competing distillation baseline by 0.94 and 0.83 percentage points and improves over pre-training without distillation by 2.31 and 2.91 points, respectively.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑