arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

权重预言机:用语言模型读取神经网络权重

Weight Oracles: Reading Neural Network Weights with Language Models

Krishna Kabra, Constantin Venhoff, Christian Schroeder de Witt

arXiv 2610.07334首次发表:更新:

发表机构

University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出权重预言机,用微调语言模型直接读取神经网络权重以诊断属性,无需行为测试,在安全审计中零样本检测后门,AUROC达0.93。

AI 中文摘要

神经网络的可解释性方法主要是反应式的:它们分析特定前向传播过程中产生的激活,需要已知输入才能发现隐藏能力(如后门)。我们提出权重预言机(Weight Oracles),即通过直接读取目标网络的原始权重来诊断其属性,而无需行为测试的微调语言模型。我们在两个阶段研究这一范式。第一阶段确立可行性:通过分阶段课程和外部计算链(将无参数操作委托给确定性代码),解释器语言模型学会从权重模拟小型变压器的前向传播,在未见过的目标上达到99%的留出准确率。第二阶段将该基础设施重新用于安全审计。我们仅使用良性病理作为训练信号,训练一个预言机来回答关于权重异常的自然语言诊断问题,并在训练中未出现的后门上对其进行零样本评估。该预言机在注意力路由后门上达到AUROC 0.93,在包括隐蔽和对抗正则化变体的多样化威胁分布上达到0.81。手工设计的统计检测器在其隐式针对的威胁模型上表现敏锐,但在威胁模型转移时崩溃,而预言机在所有攻击类型上保持均匀能力。扩展到现实规模的模型仍然是主要开放挑战。

英文摘要

Interpretability methods for neural networks are predominantly reactive: they analyse activations produced during specific forward passes, requiring known inputs to find hidden capabilities such as backdoors. We propose Weight Oracles, fine-tuned language models that diagnose properties of a target network by reading its raw weights directly, without behavioural testing. We investigate this paradigm in two phases. Phase I establishes feasibility: through a staged curriculum and an external chain-of-computation that delegates parameter-free operations to deterministic code, an explainer LLM learns to simulate the forward pass of small transformers from their weights, achieving 99% holdout accuracy on unseen targets. Phase II repurposes this infrastructure for safety auditing. We train an oracle on natural language diagnostic questions about weight anomalies using only benign pathologies as training signal, and evaluate it zero-shot on backdoors absent from training. The oracle achieves AUROC 0.93 on attention-routed backdoors and 0.81 across a diversified threat distribution including stealth and adversarially regularized variants. Hand-crafted statistical detectors are sharp on the threat models they implicitly target but collapse on threat-model shift, while the oracle remains uniformly competent across attack types. Scaling to realistic model sizes remains the principal open challenge.

CommentsSpotlight at the NeurIPS 2026 Workshop on Neural Network Artifacts as a New Data Modality (NeuralArtifacts), Paris. 14 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑