arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24482cs.LGcs.AIcs.CL

超越静态可解释性:基于预SFT参数预测后SFT机制以实现更优微调

Beyond Static Interpretability: Anticipating Post-SFT Mechanisms from Pre-SFT Parameters for Better Tuning

Hang Chen, Jiaying Zhu, Wenya Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对传统机制可解释性无法在训练前识别任务关键机制的局限,提出前瞻性定位框架,结合双粒度定位流程,实现了更优的SFT指导与稳健的可扩展性能。

中文摘要 AI 辅助

机制定位通过解释性方法分离关键参数,再以“先定位后微调”范式指导参数高效的监督微调(Supervised Fine-Tuning, SFT),从而连接机制可解释性与训练后优化。但由于机制可解释性具有回溯性,直接解释预SFT模型会引入误导性结论:针对新任务,初始识别的神经元与最终模型的控制神经元差异极大,产生的偏差会严重破坏SFT。为解决该问题,我们提出一种前瞻性定位框架,仅利用预SFT参数和目标数据集即可准确估计后SFT的可解释性状态。理论上,我们将SFT建模为连续参数演化,借助泰勒展开严格连接微调后机制目标与预SFT模型的动态梯度;实践中,我们设计了双粒度(神经元级与组件级)定位流程。大量实验表明,该方法不仅能提供更优的SFT指导,还能在模型规模增大时展现出稳健的性能与时间可扩展性。本研究突破了传统可解释性的根本局限——无法在训练前识别任务关键机制,开创了连接机制可解释性与定向优化的预测前沿。

英文摘要

Mechanistic Localization bridges mechanistic interpretability and post-training optimization by isolating critical parameters via interpretative approaches and then guiding parameter-efficient Supervised Fine-Tuning (SFT) in a ``locating-then-tuning'' paradigm. However, due to the retrospective nature of mechanistic interpretability, directly interpreting pre-SFT models introduces misleading conclusions. Specifically for novel tasks, initially identified neurons differ drastically from those governing the final model, introducing biases that actively disrupt SFT. To address this, we propose a forward-looking localization framework that accurately estimates the post-SFT interpretability state using only pre-SFT parameters and the target dataset. Theoretically, we model SFT as a continuous parameter evolution, leveraging Taylor expansion to rigorously bridge the post-tuning mechanistic objective with the pre-SFT model's dynamic gradients. Practically, we design dual-granularity (neuron- and component-level) localization pipelines. Extensive experiments demonstrate that our approach not only provides superior SFT guidance but also exhibits robust performance and temporal scalability across increasing model sizes. This work transcends the fundamental limitation of traditional interpretability-its inability to identify task-critical mechanisms before they are trained-pioneering a predictive frontier that unites mechanistic interpretability with targeted optimization.

发表机构

  • College of Computing and Data Science(计算与数据科学学院)
  • Nanyang Technological University(南洋理工大学)
  • School of Computer Science and Engineering(计算机科学与工程学院)
  • The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑