arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20991cs.LG

对齐攻击:针对图基础模型的隐蔽后门攻击

Trojaning the Alignment: Stealthy Backdoor Attacks against Graph Foundation Models

  • The Pennsylvania State University(宾夕法尼亚州立大学)
  • Amazon(亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

Minhua Lin, Zhicheng Gao, Yilong Wang, Hanqing Lu, Xiang Zhang, Suhang Wang

AI总结:

本文针对带文本属性图的图基础模型,提出STAG隐蔽特洛伊木马攻击框架,协调图触发器生成器与文本软提示实现后门攻击,经多数据集实验验证其有效性与隐蔽性。

AI中文摘要:

带文本属性图(TAGs)上的图基础模型(GFMs)将图表示与语言语义对齐,以支持可迁移的图学习。尽管具备这些优势,但TAGs上GFMs的后门漏洞仍未得到充分理解,尤其是在图-语言对齐场景中,图和文本表示被训练为在共享语义空间中相互约束。现有后门攻击主要针对图侧或文本侧,将两种模态独立对待,这使得直接适配效果不佳:仅图触发器会被干净文本语义约束,仅文本触发器会改变文本视图但不会直接改变被对齐和评分的图表示。TAGs还带来隐蔽性挑战,因为触发器同时表现为节点文本和局部图结构,使得不连贯的触发器属性或异常子图易被检测或过滤。本文提出STAG,一种针对TAGs上GFMs图-语言对齐接口的隐蔽特洛伊木马攻击框架。STAG协调图触发器生成器与文本侧软提示,使附加触发器的图表示和触发的文本表示朝向同一目标类文本区域移动。为解决TAG特定的隐蔽性问题,STAG通过候选检索将触发器节点实现为可读文本,并对附加触发器的子图进行正则化,使其局部结构与原始子图保持接近。在多个TAG数据集和代表性GFMs上的大量实验证明了STAG的有效性和隐蔽性,其代码可在该https URL获取。

英文摘要:

Graph Foundation Models (GFMs) on text-attributed graphs (TAGs) align graph representations with language semantics to support transferable graph learning. Despite these advantages, the backdoor vulnerability of GFMs on TAGs remains insufficiently understood, especially under graph-language alignment, where graph and text representations are trained to constrain each other in a shared semantic space. Existing backdoor attacks mainly target either the graph side or the text side, treating the two modalities independently. This makes direct adaptation ineffective: graph-only triggers can be constrained by clean text semantics, while text-only triggers alter the language view but do not directly shift the graph representation being aligned and scored. TAGs also impose a stealth challenge because triggers are exposed as both node text and local graph structure, making incoherent trigger attributes or anomalous subgraphs easy to inspect or filter. In this paper, we propose STAG, a stealthy trojan attack framework designed for the graph-language alignment interface of GFMs on TAGs. STAG coordinates a graph-trigger generator with a text-side soft prompt so that trigger-attached graph representations and triggered text representations move toward the same target-class text region. To address TAG-specific stealthiness, STAG realizes trigger nodes as readable text through candidate retrieval and regularizes the trigger-attached subgraph so that its local structure remains close to the original subgraph. Extensive experiments on multiple TAG datasets and representative GFMs demonstrate the effectiveness and stealthiness of STAG. Our code is available at https://github.com/ventr1c/STAG.

补充信息

↑