知晓,且仅在询问时说出:LLM 内诊学与 Minerva-7B 的精神分裂认知
Knowing, and Saying It Only When Asked: LLM Endognostics and the Schizognosis of Minerva-7B
- ISTAT(意大利国家统计局)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出LLM内诊学白盒框架,证明行为评估与内部读取系统性不一致,线性可解码性与因果控制解耦,通过消融实验验证内部知识可被因果操纵。
AI中文摘要:
通过阅读模型的回答来评估一个对齐的语言模型,这假设了回答承载着评估者所关心的区分。我们引入了 LLM 内诊学(LLM endognostics),这是一个白盒内部审计框架,旨在提取并因果性地操纵残差流中的潜在知识。将该框架应用于 Minerva-7B-Instruct-v1.0,在12类职业风险上的124个最小提示对中,行为评估在大多数集合上失败:模型在63.7%的提示对上表现相同(95%置信区间[55.0%, 71.6%]),即同时遵从或拒绝两个成员。然而,通过雅可比透镜将残差流投影到词汇表上,揭示了一个统计上显著的对比性内诊边际(Contrastive Endognostic Margin),证明模型在内部保持了稳健的风险区分。在第二个协议中,交叉了25个事实与五种语言框架,我们显示模型在72%的情况下遵从了预设的虚假信息,尽管在其潜在层中表征了真实实体。外科手术式地消融植入虚假信息的方向,在25个被抑制案例中的11个中恢复了正确答案(McNemar检验p = 0.0010),并通过盲人人工标注验证(二元一致性kappa = 0.68)。相比之下,一个样本外线性探针在第10层达到77%的准确率,但其正交消融的恢复率为0%。这建立了一个基本的理论分离:抽象的线性表征并不意味对言语化的因果控制。我们的主要贡献是形式化了内诊学,以证明行为评估和内部读取在常见情况下系统性不一致,并且线性可解码性与因果控制是解耦的。
英文摘要:
Evaluating an aligned language model by reading its answers assumes the answers carry the distinction the evaluator cares about. We introduce LLM endognostics, a white-box internal auditing framework designed to extract and causally manipulate latent knowledge within the residual stream. Applied to Minerva-7B-Instruct-v1.0 on 124 minimal prompt pairs over 12 categories of professional risk, behavioral evaluation fails on most of the set: the model acts identically on 63.7% of the pairs (95% CI [55.0%, 71.6%]), complying with or refusing both members. Yet, projecting the residual stream onto the vocabulary by a Jacobian lens reveals a statistically significant Contrastive Endognostic Margin, proving the model maintains robust risk differentiation internally. In a second protocol crossing 25 facts with five linguistic framings, we show that the model conforms to presupposed falsehoods in 72% of cases, despite representing the true entity in its latent layers. Surgically ablating the direction of the planted falsehood restores the correct answer in 11 of 25 suppressed cases (McNemar p = 0.0010), validated by blind human annotation (binary agreement kappa = 0.68). In contrast, an out-of-sample linear probe achieves 77% accuracy at layer 10, but its orthogonal ablation yields a 0% recovery rate. This establishes a fundamental theoretical dissociation: abstract linear representation does not imply causal control over verbalization. Our main contribution is the formalization of endognostics to prove that behavioral evaluation and internal reading systematically disagree in the common case, and that linear decodability is decoupled from causal control.