arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

路由子空间:审计微调语言模型中评估到部署的不匹配

Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models

Phongsakon Mark Konrad, Toygar Tanyel, Serkan Ayvaz

arXiv 2607.20436首次发表:更新:

发表机构

Centre for Industrial Software, University of Southern Denmark; ProMake(南丹麦大学工业软件中心; 宝迈可)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究微调语言模型评估到部署的不匹配问题,提出在路径修补告知的中深度窗口拟合配对激活对比并修改结果坐标的方法,在多个模型实例中缩小了差距,还指出单坐标审计的局限性,此审计用于诊断微调检查点。

AI 中文摘要

安全评估通常假设测试期间观察到的行为反映了实际使用中的行为,但微调可能会打破这一假设。一个检查点在评估式提示下看似固定,而在实际使用提示下却有不同行为。输出分数能揭示这种不匹配但无法定位。我们研究这种差异是否编码在稳定的内部位置,并引入一种方法,在路径修补告知的中深度窗口拟合配对激活对比,然后在保留提示上修改结果坐标。该干预在四个全矩阵指令微调模型实例的十二个模型 - 行为设置中的十个中缩小了评估到部署的差距;第五个模型支持定位和编辑来源检查,部署框架率变化最多6.1个百分点。两个扁平单元表明,当安装的差异是高阶或被深度启发式错过时,单坐标审计是不够的。该审计是对微调检查点的诊断,而非训练时的防御或部署安全的保证。

英文摘要

Safety evaluations often assume that behavior observed during testing reflects behavior in ordinary use, but fine-tuning can break this assumption. A checkpoint can appear fixed under evaluation-style prompts while the same behavior persists under ordinary-use prompts. Output scores reveal this mismatch but do not locate it. We investigate whether the distinction is encoded in a stable internal site and introduce an approach that fits a paired activation contrast at a path-patching-informed mid-depth window, then modifies the resulting coordinate on held-out prompts. The intervention closes the evaluation-to-deployment gap in ten of twelve model--behavior settings (six of the eight settings with $n{\geq}120$ paired questions) across four full-matrix instruction-tuned model instances; a fifth model supports localization and edit-provenance checks, and deployment-framed rates change by at most $6.1$pp. The two flat cells, both sycophancy, indicate that a single-coordinate audit is not sufficient when the installed distinction is higher-rank or missed by the depth heuristic. The audit is a diagnostic for fine-tuned checkpoints, not a training-time defense or a guarantee of deployment safety.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑