arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从方向到大小:多模态指令微调如何在Transformer隐藏状态中重新组织身份指定提示的几何编码

System-Prompt Conditioning and Hidden-State Geometry in Four Open-Weight Models: Corrections and What Survives

Jorge Castillo Sepúlveda, Marco Torres Yévenes, Juan Carlos Lanas

arXiv 2607.09842首次发表:更新:

发表机构

Axis Dynamics SpA(轴动力公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究四种训练后模式的Transformer语言模型,通过五个几何指标比较三种提示条件,发现多模态指令微调会使身份编码从方向重组为大小,此重组特定于该模式,还定位了W_1作为方法贡献。

AI 中文摘要

我们研究了身份指定系统提示是否会在跨越四种训练后模式的四个开放权重Transformer语言模型的隐藏状态轨迹中产生统计学上可区分的几何指纹:无训练(Gemma-4-E4B基础模型)、多模态基于人类反馈的强化学习(Gemma-4-E4B-it)、强化学习蒸馏(DeepSeek-R1-Distill-Qwen-7B)和监督微调(Qwen2.5-7B-Instruct)。通过五个几何指标比较了三种提示条件(身份指定轴提示、长度匹配的通用助手提示和26-token的普通基线),主要是k近邻轨迹图上Ollivier-Ricci曲率的边分布之间的1-Wasserstein距离。研究结果基于具有多个几何控制的轨迹级排列测试。核心发现是跨指令微调边界的身份编码有定性重组:在基础模型中指纹是方向编码的(角度k近邻下分离度为0.034,p = 0.002);在多模态指令微调模型中它迁移到大小上(角度分离度降至p = 0.439,而欧几里得距离在p = 0.042时存在,并且第一个生成状态的平均范数反转了其长度顺序,身份提示的最低)。这种从方向到大小的重组特定于多模态指令微调模式,在强化学习蒸馏和监督微调下不存在。教师强制控制将约30%的自由运行余弦信号归因于提示驱动的效果。我们将k近邻轨迹图上边的Ollivier-Ricci分布上的W_1定位为具有独立研究价值的方法贡献。

英文摘要

Versions 1 and 2 of this preprint reported that an identity-specifying system prompt leaves a geometric fingerprint in the final-layer hidden-state trajectories of four open-weight language models, and that instruction tuning moves this fingerprint from the direction to the magnitude of the hidden-state vector. An audit of their code and data found the following. The curvature statistic described as Ollivier-Ricci curvature on Euclidean k-NN graphs was a non-standard Forman-type edge statistic on graphs built from temporal and cosine k-NN edges. Its released test permuted pooled edges instead of trajectories, and the published p-values came from unreleased code. The quantity reported as the norm of the first generated state is the state at the last prompt position, from which the first output token is predicted. The generic control prompt was matched to the identity prompt in characters, not in tokens. This version corrects the methods, withdraws the regime-specific claims (one model per regime) and the direction-to-magnitude claim, and re-analyzes the data with added controls. What survives is narrower. Centroid distance, maximum mean discrepancy and a linear probe separate every pair of prompt conditions in every model, while the curvature statistic exceeds its split-half noise floor in only four of twelve comparisons. In Gemma-4-E4B-it this state has a lower norm under the identity prompt than under a token-length-matched generic prompt (138.1 vs. 216.5; Cohen's d = -5.45; n = 20). Its direction also separates the conditions, and the effect fits the state's role in planning the output: the identity prompt instructs a pause before every answer, and the model opens 98 of 100 responses with a pause marker. When the first token is fixed, the norm ordering reverses. The base model continues the prompt template instead of answering. A redesigned follow-up study is in preparation.

Commentsv3: substantial correction, replaces v1-v2. The curvature reported as Ollivier-Ricci was a Forman-type statistic on temporal plus cosine k-NN graphs, tested at edge level. The regime-specific and direction-to-magnitude claims are withdrawn. Added: simple-baseline, split-half, trajectory-level and token-length-matched controls. 14 pages, 1 figure, 9 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑