arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

跨语言模型家族预测训练响应的目标无关微干预

Target-Independent Micro-Interventions for Predicting Training Response Across Language-Model Families

Zhongxuan Liu, Sicheng Zhou, Hongzhi Wang

arXiv 2609.08618首次发表:更新:

发表机构

Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出目标无关微干预方法,通过分支标准化干预构建L-State,利用直接和算子读出预测跨家族语言模型的训练响应,显著降低MSE并提升符号平衡准确率。

AI 中文摘要

基准分数描述了检查点当前能做什么,但并不能决定它在下一次训练阶段中将如何响应。我们通过从同一检查点分支四个简短、标准化、目标无关的微干预,并在一个共同的能力空间中记录其效果,来衡量这一缺失的状态。结合当前能力,这些响应构成了L-State;其脉冲块支持灵活的直接读出和保持结构的算子读出。在平滑局部动力学下,算子构造允许端到端的跨家族界限,并显式考虑源家族和目标家族的坐标异质性。在三家族留一家族开发中,两种脉冲读出相对于仅使用能力,将源标准化均方误差降低了39.4%,同时分离了最佳响应和方向估计。在密封的GLM-4-9B上,直接读出和算子读出分别将均方误差降低了71.8%和78.3%,算子读出将符号平衡准确率从0.366提高到0.754。在密封的Granite-3.1-8B上,直接读出达到均方根误差0.544,开发拟合的动作选择器达到0.554,而仅使用能力时为1.172。一项五家族审计发现,算子坐标随动作和家族而变化,并且对这些偏差进行建模可改善回顾性保持轨迹预测。因此,目标无关干预暴露了当前能力所遗漏的训练响应信息,直接读出和结构化读出覆盖了互补的迁移机制。

英文摘要

Benchmark scores describe what a checkpoint can do now, but they do not determine how it will respond to the next training episode. We measure this missing state by branching four short, standardized, target-independent micro-interventions from the same checkpoint and recording their effects in a common capability space. Together with current capability, these responses form L-State; its pulse block supports a flexible direct readout and a structure-preserving operator readout. Under smooth local dynamics, the operator construction admits an end-to-end cross-family bound with explicit source- and target-family coordinate heterogeneity. In three-family leave-one-family-out development, both pulse readouts reduce source-standardized MSE by 39.4% relative to capability alone, while separating the best response and direction estimates. On sealed GLM-4-9B, the direct and operator readouts reduce MSE by 71.8% and 78.3%, respectively, and the operator readout raises sign balanced accuracy from 0.366 to 0.754. On sealed Granite-3.1-8B, the direct readout reaches RMSE 0.544 and a development-fitted action-wise selector reaches 0.554, compared with 1.172 for capability alone. A five-family audit finds that the operator coordinate varies by action and family, and that modeling these deviations improves retrospective held-trajectory prediction. Target-independent interventions therefore expose training-response information that current capability misses, with direct and structured readouts covering complementary transfer regimes.

Comments24 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑