AI 中文总结
研究通过攻击图基础模型的共享表示来探索其脆弱性,采用表示空间扰动和输入空间攻击等方法,针对六个公共模型进行实验,发现不同模型的脆弱性表现,如OpenGraph的频谱分词器有特定脆弱性,还分析了脆弱性与解码器读取表示的关系。
AI 中文摘要
图基础模型通过在任何任务推理之前将每个输入映射到一个共享表示来跨图域进行泛化。我们将此映射称为对齐层,它是将图基础模型与图神经网络区分开来的组件,并且我们表明它是一个先前工作未研究过的独特攻击面。我们在推理时对其进行攻击,在不接触训练的情况下,针对跨越频谱分词器、文本嵌入空间和离散码本的六个公共模型进行攻击。一种有向表示空间扰动会使每个模型崩溃,但所需预算与普通图网络所需的表示范数相当,只有一个例外:OpenGraph,其频谱分词器在五分之一的预算下崩溃,这是普通网络不具备的特定于对齐的脆弱性,并且相同表示控制将其追溯到分词器而不是解码器。一种可实现的输入空间攻击,即编辑边、特征或文本,在峰值时会使六个模型中的三个模型至少一半的正确预测被移除。输入访问攻击者实现的这种脆弱性程度与解码器读取表示的直接程度有关,而不是与任务的干净准确率有关;我们从解码器的局部利普希茨敏感性在结构上测量这种载体增益,并将干净准确率余量报告为模型内的排序启发式方法,这种方法在可实现的攻击中无法存活。
英文摘要
A graph foundation model generalizes across graph domains by mapping every input into one shared representation before any task reasoning. We call this map the alignment layer, the component that separates a graph foundation model from a graph neural network, and we show it is a distinct attack surface that prior work has not studied. We attack it at inference time, with no access to training, on six public models spanning spectral tokenizers, text embedding spaces, and a discrete codebook. A directed representation-space perturbation collapses every model, but at a budget comparable to the representation norm a plain graph network also needs, with one exception: OpenGraph, whose spectral tokenizer collapses at a fifth of that budget, an alignment-specific fragility a plain network does not share and which a same-representation control traces to the tokenizer rather than the decoder. A realizable input-space attack that edits edges, features, or text removes at least half the correct predictions on three of the six models at peak. How much of this fragility an input-access attacker realizes tracks how directly the decoder reads the representation, and not the clean accuracy a task leaves; we measure this carrier gain structurally from the decoder's local Lipschitz sensitivity, and report clean-accuracy headroom as a within-model ordering heuristic that does not survive on realizable attacks.