发表机构
East China Normal University(华东师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出TRACE方法,利用检查点更新中的几何痕迹,无需模型查询即可审计有害SFT,在多种设置下稳定且与攻击成功率正相关。
AI 中文摘要
对训练后大型语言模型的安全审计通常依赖于模型行为,这需要执行模型并依赖于可用评估的覆盖范围。本研究提出了一个不同的问题:在监督微调(SFT)过程中优化的目标行为是否会在检查点更新中直接留下可读的证据?我们发现,有害顺从性SFT在检查点更新空间中诱导了一个连续的、目标依赖的排序。利用由纯有害顺从性、安全定向和良性效用SFT定义的参考几何结构,我们发现检查点级别的坐标s_H在四个7-8B骨干模型上以0.986-0.992的Spearman相关性跟踪受控的有害目标组成,且相同的排序在更大模型规模上持续存在。匹配的顺从性对比拒绝性对照表明,这种检查点痕迹反映了SFT目标而非有害输入暴露,而额外的对照排除了基于有害示例数量或通用训练强度的简单解释。基于这一结构,我们引入了TRACE,一种仅权重的审计方法,它将未知的检查点更新相对于冻结的有害和非有害参考原型进行定位,并将这种几何结构转换为连续的有害目标分数。TRACE既不需要模型查询,也不需要访问未知的SFT数据,可以直接从检查点更新中评估。在分布偏移、未见数据、不同的SFT配置、部分检查点访问以及LoRA/全参数微调下,痕迹保持稳定,并与独立测量的攻击成功率正相关。即使在低有害目标比例下,TRACE仍然具有信息量,在行为评估不可用或不完整时提供补充的审计信号。代码可在以下URL获取。
英文摘要
Safety auditing of post-trained large language models typically relies on model behavior, requiring model execution and depending on the coverage of available evaluations. This work asks a different question: Do the target behaviors optimized during supervised fine-tuning (SFT) leave readable evidence directly in checkpoint updates? We find that harmful-compliance SFT induces a continuous, objective-dependent ordering in checkpoint-update space. Using a reference geometry defined by pure harmful-compliance, safety-targeted, and benign-utility SFT, we find that a checkpoint-level coordinate s_H tracks controlled harmful-objective composition with Spearman correlations of 0.986-0.992 across four 7-8B backbones, with the same ordering persisting at larger model scales. Matched compliance-versus-refusal controls show that this checkpoint trace reflects the SFT objective rather than harmful-input exposure, while additional controls rule out simple explanations based on harmful-example count or generic training intensity. Building on this structure, we introduce TRACE, a weights-only auditing method that localizes an unknown checkpoint update relative to frozen harmful and non-harmful reference prototypes and converts this geometry into a continuous harmful-objective score. TRACE requires neither model queries nor access to the unknown SFT data, and can be evaluated directly from checkpoint updates. Across distribution shifts, unseen data, different SFT configurations, partial checkpoint access, and LoRA/full-parameter fine-tuning, the trace remains stable and is positively associated with independently measured attack success rates. TRACE remains informative even at low harmful-objective proportions, providing a complementary auditing signal when behavioral evaluation is unavailable or incomplete. Code is available at https://anonymous.4open.science/r/Code4TRACE-54D3.