先回答再编辑:用于保留效用的反蒸馏的推理骨架编辑
Answer-then-Edit: Reasoning Skeleton Editing for Anti-Distillation with Preserved Utility
- School of Computer Science and Engineering, UNSW Sydney(新南威尔士大学计算机科学与工程学院)
- School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对专有大语言模型易受知识蒸馏影响的问题,提出骨架引导推理编辑框架SGRE,通过先回答再编辑的方式,借鉴认知负荷理论对推理痕迹进行修改,在降低蒸馏效果的同时保持推理准确性和痕迹自然度。
AI中文摘要:
专有大语言模型需要大量智力和资金投入,使其成为有价值的知识产权。即便通过黑盒API部署,这些模型仍易受未经授权的知识蒸馏影响。为解决此问题,反蒸馏被提出以生成阻碍蒸馏效果的防御输出。然而,现有基于内部模型扰动的反蒸馏方法难以平衡推理痕迹的抗蒸馏性和效用。为此,我们提出了骨架引导推理编辑(SGRE),这是一个用于反蒸馏的事后痕迹修改框架。在回答阶段,教师模型生成干净的推理痕迹,保留原始推理准确性并灵活控制痕迹自然度。在编辑阶段,借鉴认知负荷理论,引入由推理骨架提取、骨架图粗化和骨架语言化组成的三阶段策略。大量实验表明,SGRE在降低蒸馏效果方面达到了先进性能,同时保持无损推理准确性和卓越的痕迹自然度。
英文摘要:
Proprietary large language models (LLMs) entail substantial intellectual and financial investment, making them valuable intellectual property (IP). However, even when deployed via black-box APIs, these models remain vulnerable to unauthorized knowledge distillation, which allows adversaries to cheaply extract and replicate model capabilities. To address this issue, anti-distillation (AD) has been proposed to generate defensive outputs that hinder distillation effectiveness, overcoming the limitation of watermarking-based approaches that rely on post-hoc verification. However, existing AD methods based on internal model perturbations struggle to balance anti-distillability and utility (e.g., answer accuracy and naturalness) of reasoning traces, with stronger defenses often causing significant utility loss. To fill this gap, we propose \textbf{\underline{S}}keleton-\textbf{\underline{G}}uided \textbf{\underline{R}}easoning \textbf{\underline{E}}diting (SGRE), an \textit{Answer-then-Edit} framework that performs post-hoc trace modification for anti-distillation. In the answer stage, the teacher model first generates clean reasoning traces, preserving the original reasoning accuracy while enabling more flexible control over trace naturalness. In the editing stage, we draw inspiration from Cognitive Load Theory (CLT) and introduce a three-stage strategy consisting of reasoning skeleton extraction, skeleton graph coarsening, and skeleton verbalization. These operations jointly perturb reasoning structures and augment textual complexity to amplify extraneous load on student models, hindering their acquisition of underlying reasoning patterns. Extensive experiments across diverse LLMs demonstrate that SGRE achieves state-of-the-art performance in reducing distillation effectiveness, while maintaining lossless reasoning accuracy and superior trace naturalness.