arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08572cond-mat.stat-mech

AI知识蒸馏中的可靠性-安全性权衡:一种重整化群方法

Reliability-Safety Trade-off in AI Distillation: A Renormalization-Group Approach

Y. M. Du, Miao-Miao Yi, Tan-Ji Zhou, C. P. Sun

首次发表
浏览论文内容

中文总结 AI 辅助

该研究基于统计力学构建知识蒸馏的粗粒化模型,以参数K表征危险辨别能力,揭示了AI蒸馏中可靠性与安全性的权衡关系,为多代蒸馏提供了可检验的标度预测。

中文摘要 AI 辅助

知识蒸馏传递的不仅是任务能力,还包括响应倾向、拒绝策略、错误边界以及潜在的安全偏差。我们将这种行为继承建模为基于统计力学的粗粒化模型,其中学生模型的答案与弃权(不执行)决策定义了两个宏观态,而教师模型诱导的有效场会重塑学生模型的自由能景观。该模型得出由单一参数K(我们称之为危险辨别能力)控制的可靠性-安全性权衡关系,预测的权衡与弃权标记数据[arXiv:2412.06748]一致。在知识蒸馏中,具有强危险辨别能力的教师模型可提升学生模型可达的可靠性与安全性,而辨别能力差则会限制可达的权衡。重复蒸馏相当于迭代的类重整化群变换,K会在多代间发生流动,该流动呈现出三临界结构,将K损失、稳定传递及阈值依赖继承的区域分隔开,并为多代蒸馏提供可检验的标度预测。

英文摘要

Knowledge distillation transfers more than task competence: it also transmits response propensities, refusal policies, error boundaries, and latent safety biases. We formulate this behavioral inheritance as a coarse-graining model grounded in statistical mechanics, in which the student's answer and refusal decisions define two macrostates, while the teacher induces an effective field that reshapes the student's free-energy landscape. The model yields a reliability-safety trade-off relation controlled by a single parameter K, which we term the hazard discrimination capability. The predicted trade-off is consistent with refusal-token data [arXiv: 2412.06748]. In knowledge distillation, a teacher with strong hazard discrimination improves the student's attainable reliability and safety, whereas poor discrimination limits the attainable trade-off. Repeated distillation acts as an iterated renormalization-group-like transformation, under which K follows a flow across generations. The flow exhibits a tricritical structure separating regimes of K loss, stable transmission, and threshold-dependent inheritance, and yields testable scaling predictions for multigenerational distillation.

↑