arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自我教师应该看到什么?在策略自蒸馏中的特权上下文设计

What Should a Self-Teacher See? Privileged Context Design for On-Policy Self-Distillation

Kanghui Tian, Siyuan Liu, Tianxiang Jiang, Shuai Dong, Yizhuo Li, Tian Ding, Yuan Guo, Songze Li, Haowen Hou, Congcong Wang, Yi Wang

arXiv 2609.25623首次发表:更新:

发表机构

Fudan University; Shanghai Artificial Intelligence Laboratory; Nanjing University; Shanghai Jiao Tong University; Peking University; University of California, Los Angeles; Tongji University(复旦大学; 上海人工智能实验室; 南京大学; 上海交通大学; 北京大学; 加利福尼亚大学洛杉矶分校; 同济大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探讨在策略自蒸馏中特权上下文的设计,发现中间抽象级别(如命名策略、框架或类别)优于完整解决方案,能提升数学任务性能并减少提示存储,最佳上下文取决于学生规模和任务。

AI 中文摘要

更多的特权信息并不总是造就更好的教师。我们在在策略自蒸馏(OPSD)中研究这一张力,其中基础模型的冻结副本在特权上下文下对学生的自身轨迹进行评分,特权上下文通常是一个完整的参考解决方案,将最终答案与一条特定的推理路径捆绑在一起。在每个规模内保持学生视图和训练固定,我们将该默认设置与三种离线编译的抽象进行比较:一种命名策略、一种方法无关的框架和一种问题类别,并与一个仅答案的对照进行比较,该对照保留目的地但移除路径。在竞赛数学的主要运行中,最佳中间上下文在4B规模上将域内峰值均值比完整解决方案提高1.4个点,在8B规模上提高1.6个点,同时存储的提示令牌少一个数量级。跨三个种子的比较也显示,在两种规模下,框架和类别上下文均获得正的平均增益。仅答案的条件在主要运行中保持竞争力,在这些规模下与完整解决方案的差距在0.2个点以内。首选上下文随学生规模和任务而变化。初始师生KL散度不对下游性能进行排序。因此,自我教师应该看到的不是它所能看到的一切,而是其学生仍能据此行动的那种抽象层次。

英文摘要

More privileged information does not always make a better teacher. We study this tension in on-policy self-distillation (OPSD), where a self-teacher scores the student's own rollouts under privileged context, conventionally a complete reference solution that bundles the final answer with one particular reasoning path. Holding the student view and training fixed within each scale, we compare that default against three abstractions compiled offline, a named strategy, a method-independent framing, and a problem category, and against an answer-only control that keeps the destination but removes the path. In the primary runs on competition mathematics, the best intermediate contexts improve the in-domain peak mean over the full solution by 1.4 points at 4B and 1.6 at 8B, while storing an order of magnitude fewer hint tokens. Comparisons across three seeds also show positive mean gains for the framing and category contexts at both scales. Answer-only conditioning remains competitive in the primary runs, within 0.2 points of the full solution at these scales. The preferred context varies with student scale and task. Initial teacher-student KL does not order downstream performance. What a self-teacher should see is therefore not everything it could, but the level of abstraction its student can still act on.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑