发表机构
Scuola Superiore Sant'Anna(圣安娜高等学校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对ViT拆分推理的Token缩减与打乱机制,提出SARA攻击方法,发现Token打乱仅具表面隐私、Token缩减仍存泄露,还提出轻量边缘侧防御以降低攻击性能且无需改动云端模型。
AI 中文摘要
视觉Transformer(ViT)越来越多地应用于拆分推理系统,在此类系统中,边缘设备会将中间Token表示传输至远程云端。在该场景下,Token缩减可降低计算与通信成本,而Token打乱会破坏所传输Token的空间结构,可能限制信息泄露。然而,它们针对特征逆攻击的隐私保护效果仍不明确,特征逆攻击旨在从传输的嵌入中重建输入。本研究表明,尽管Token打乱破坏了传统重建攻击所需的空间结构,但传输的Token嵌入仍保留大量位置信息。基于此观察,我们提出空间对齐重建攻击(Spatially Aligned Reconstruction Attack,SARA),这一统一流程可预测Token位置、恢复其空间布局、使用特征空间掩码自动编码器重建缺失嵌入,并恢复输入图像。我们的结果显示,Token打乱仅提供表面隐私,因为SARA在很大程度上重建了原始Token结构;Token缩减提供更强保护,但当保留的Token具备足够语义与位置信息时,仍存在显著泄露。最后,我们提出一种轻量级边缘侧防御方法,该方法会移除位置嵌入并通过知识蒸馏逐步调整边缘侧Transformer模块,其可大幅降低针对SARA的攻击性能,同时保留下游任务精度且无需修改云端模型。
英文摘要
Vision Transformers (ViTs) are increasingly used in split-inference systems, where edge devices transmit intermediate token representations to a remote cloud. In this setting, token reduction lowers computation and communication costs, while token shuffling disrupts the spatial organization of the transmitted tokens, potentially limiting information leakage. However, their privacy benefits remain unclear against feature inversion attacks, which attempt to reconstruct the input from the transmitted embeddings. In this work, we show that, despite disrupting the spatial structure required by conventional reconstruction attacks, transmitted token embeddings retain substantial positional information. Based on this observation, we introduce the Spatially Aligned Reconstruction Attack (SARA), a unified pipeline that predicts token positions, restores their spatial layout, reconstructs missing embeddings using a feature-space masked autoencoder, and recovers the input image. Our results demonstrate that token shuffling provides only apparent privacy, as SARA largely reconstructs the original token organization. Token reduction offers stronger protection, but significant leakage persists when the retained tokens preserve sufficient semantic and positional information. Finally, we introduce a lightweight edge-side defense that removes positional embeddings and progressively adapts the edge-side transformer blocks through knowledge distillation. It substantially reduces attack performance against SARA, while preserving downstream task accuracy and requiring no changes to the cloud-side model.
CommentsAccepted at the 19th ACM Workshop on Artificial Intelligence and Security (AISec'26)