arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不要信任AI生态系统:分析受感染开源组件中的隐私泄露

Don't Trust the AI Ecosystem: Analyzing Privacy Leakage in Compromised Open-Source Components

Jin-Seong Kim, Han-Ju Lee, Seok-Won Hong, Takeshi Takahashi, Chansu Han, Tomohiro Morikawa, Seok-Hwan Choi

arXiv 2607.27886首次发表:更新:

AI 中文总结

该研究提出GradLock新型训练时注入攻击,将敏感数据注入模型参数,可从最终模型提取像素级完美数据,鲁棒性强且难被检测,揭示AI供应链安全存在严重盲点。

AI 中文摘要

现有的模型反演(MI)攻击主要依赖训练后优化,从模型输出中恢复隐私数据,但这些方法受目标模型泛化瓶颈的根本限制,常产生通用特征而非特定身份,高维数据集上尤为明显。本文提出GradLock,一种新型训练时注入攻击,将敏感训练数据隐秘注入模型参数,在受感染供应链场景下,该攻击利用无状态确定性索引建立隔离数据 vault,采用动态梯度锁定防止优化过程中 payload 退化,使攻击者无需访问训练环境即可从最终模型中提取像素级完美数据。在MNIST、Imagenette、CelebA上的大量实验表明,GradLock实现近无损重建(SSIM≈1.0)且提取耗时<1.0秒;与现有训练时注入方法相比,其对量化、剪枝、微调等标准部署优化的鲁棒性更优。此外,用户部署研究显示93.3%的参与者未检测到恶意逻辑,凸显现代AI供应链安全存在严重盲点。

英文摘要

Existing model inversion (MI) attacks predominantly rely on post-training optimization to recover private data from model outputs. However, these methods are fundamentally constrained by the target model's generalization bottleneck, often yielding generic features rather than specific identities, particularly on high-dimensional datasets. In this paper, we introduce GradLock, a novel training-time injection attack that stealthily injects sensitive training data directly into the model parameters. Operating within a compromised supply chain context, GradLock leverages stateless deterministic indexing to establish isolated data vaults and employs dynamic gradient locking to prevent payload degradation during the optimization process. This mechanism allows the adversary to extract pixel-perfect data from the final model without retaining access to the training environment. Extensive experiments on MNIST, Imagenette, and CelebA demonstrate that GradLock achieves near-lossless reconstruction (SSIM ~ 1.0) and instant extraction (< 1.0s). Compared to existing training-time injection methods, our approach exhibits superior robustness against standard deployment optimizations, including quantization, pruning, and fine-tuning. Furthermore, a user deployment study reveals that 93.3% of participants failed to detect the malicious logic, highlighting a severe blind spot in the security of modern AI supply chains.

CommentsExtended version of the paper to appear in ACM CCS 2026 (17 pages, 10 figures)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑