arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25727cs.LGcs.CR

增强型大型语言模型的图神经网络是否具备隐私安全性?

Auditing Privacy Risks in LLM-Enhanced Graph Neural Networks

发表机构北京邮电大学
查看机构详情
  • Beijing University of Posts and Telecommunications(北京邮电大学)

机构由 AI 辅助整理,请以论文原文为准。

Longzhu He, Zelang Wen, Chaozhuo Li, Sen Su

首次发表
浏览论文内容

中文总结 AI 辅助

本文通过5阶段统一框架评估LLM增强型GNNs的隐私风险,发现其隐私脆弱性高于浅层文本基线,差分隐私可部分缓解风险但会导致效用下降,凸显了该领域的隐私-效用权衡。

中文摘要 AI 辅助

大型语言模型(LLMs)近期通过用语义信息丰富节点表示,推动了图神经网络(GNNs)的发展,催生的LLM增强型GNNs取得了显著的性能提升。然而,其对隐私攻击的脆弱性——即攻击者可从模型输出中推断敏感信息——在很大程度上仍未被探索。为填补这一空白,本文通过一个包含五个阶段的统一框架,对LLM增强型GNNs的隐私风险开展系统性评估:(1)数据集准备;(2)受害模型训练;(3)隐私攻击;(4)风险评估;(5)防御分析。具体而言,本文在覆盖不同领域的6个真实世界文本属性图数据集上开展实验,考虑了针对链路、标签和成员推断这三类基础威胁的6种代表性隐私攻击方法,通过将多种基于LLM的特征增强器与代表性GNN主干相结合,构建了42种受害模型配置。大量实验表明,尽管LLM增强型GNNs的效用有所提升,但与浅层文本表示基线相比,其对隐私攻击的脆弱性始终更高。进一步分析显示,语义丰富会放大嵌入空间中与链路、标签和成员相关的信号,使其更易被推断攻击利用。最后,本文评估了差分隐私作为防御策略的效果,结果表明,尽管差分隐私可部分缓解隐私风险,但会导致显著的效用下降,凸显了LLM增强型图学习中存在的根本性隐私-效用权衡。总体而言,本研究为全面理解LLM增强型GNNs的隐私风险提供了支撑,并为开发更安全、可信赖的图学习系统提供了实用见解。

英文摘要

Large language models (LLMs) have recently advanced graph neural networks (GNNs) by enriching node representations with semantic information, giving rise to LLM-enhanced GNNs that achieve substantial performance gains. However, how such semantic enhancement affects privacy risks remains largely underexplored. To bridge this gap, we systematically audit the privacy risks of LLM-enhanced GNNs through a unified framework consisting of five stages: (1) dataset preparation, (2) victim model training, (3) privacy attack, (4) risk assessment, and (5) defense analysis. Specifically, our evaluation spans ten text-attributed graph datasets across diverse domains, six privacy attacks, 42 LLM-enhanced GNN configurations, and three more recent language-model backbones. Extensive experiments show that, despite their utility improvements, LLM-enhanced GNNs consistently exhibit greater empirical privacy vulnerability than shallow text representation baselines under the evaluated attacks across diverse models and datasets. Further analysis shows that LLM-enhanced representations exhibit more distinguishable link-, label-, and membership-related signals in the embedding space, making them more exploitable by inference attacks. Finally, we evaluate representative defenses and examine their effectiveness in mitigating these privacy risks. Overall, this work provides a systematic audit of privacy risks in LLM-enhanced GNNs and offers insights for developing more secure and trustworthy graph learning systems.

↑