arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向大语言模型的通用文本教学

Universal Textual Teaching for LLMs

Zhanyi Lu, Huan Wang

arXiv 2610.12114首次发表:更新:

发表机构

Westlake University(西湖大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出无需参数更新的通用文本教学(UTT)框架,通过多角色交互生成可复用的入门指南,在数学、代码生成任务上显著提升学生模型性能,且可跨模型泛化。

AI 中文摘要

知识蒸馏(KD)将知识从更强的教师模型传递到更弱的学生模型,但大多数方法需要训练学生模型的参数,从而将蒸馏得到的知识绑定到特定的架构和检查点。这种隐式表示难以解释,也难以在不同模型间复用,限制了KD在仅提供API或训练成本高昂的模型中的应用。本文研究大语言模型(LLM)的知识传递,提出通用文本教学(UTT),这是一种无需参数更新的框架,可将观察到的教师-学生知识差距蒸馏为一种文本化、可解释且可复用的自然语言产物,称为“入门指南(Primer)”。具体而言,UTT首先通过配对评估识别代表性差距案例,再通过多角色交互迭代更新入门指南:学生尝试完成每个任务,提示器将评估反馈转化为教学指令,教师提供针对性演示,合成器整合已验证的经验。实验方面,在具有挑战性的数学任务(Omni-MATH-2)和代码生成任务(KernelBench)上,大量结果证实了该方法的有效性:UTT将KernelBench上学生模型的准确率从9.4%显著提升至48.6%,Fast1准确率从9%提升至35%,同时将数学推理准确率从27.6%提升至51.7%。UTT的性能也优于代表性的提示工程和基于参数的KD方法。值得注意的是,UTT被证明可在不同教师和学生间泛化:为某一教师-学生对合成的入门指南可泛化到未参与合成的其他学生模型。

英文摘要

Knowledge distillation (KD) transfers knowledge from stronger Teacher models to weaker Student models, but most methods require training the Student parameters, thereby binding the distilled knowledge to a specific architecture and checkpoint. This implicit representation is difficult to interpret or reuse across models and limits KD for API-only or costly-to-train models. This paper studies knowledge transfer for large language models (LLMs). We introduce Universal Textual Teaching (UTT), a parameter-update-free framework that distills observed Teacher-Student knowledge gaps into a textual, interpretable, and reusable natural-language artifact called Primer. Specifically, UTT first identifies representative gap cases through paired evaluations, and iteratively updates the Primer via multi-role interactions: the Student attempts each task, the Prompter turns evaluation feedback into a teaching instruction, the Teacher provides a targeted demonstration, and the Synthesizer consolidates validated lessons. Empirically, on the challenging math (Omni-MATH-2) and code generation (KernelBench) tasks, extensive results confirm the effectiveness of the method: UTT remarkably raises the Student's accuracy from 9.4% to 48.6% and Fast1 accuracy from 9% to 35% on KernelBench, while increasing mathematical reasoning accuracy from 27.6% to 51.7%. UTT also performs better than representative prompt engineering and parameter-based KD methods. Of note, UTT is shown to be generalizable across different Teachers and Students: a Primer synthesized for one Teacher-Student pair can generalize to other Students that do not participate in the synthesis.

CommentsWebpage: https://alexlu99.github.io/UTT/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑