arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15671cs.CVcs.AIcs.CR

不要发送不需要的内容:面向视觉语言模型的基于问题引导的令牌剪枝隐私防御

Don't Send What You Don't Need: Question-Guided Token Pruning as a Privacy Defense for Vision-Language Models

  • Southern Illinois University Carbondale(南伊利诺伊大学卡本代尔分校)

机构由 AI 辅助整理,请以论文原文为准。

Md Khalid Syfullah, Alvi Ataur Khalil

AI总结:

提出QPriv-VL框架,利用动态阈值预测器基于问题相关性剪枝视觉令牌,在联邦/分割学习中防御隐私攻击,以约40%令牌保持准确率并显著降低攻击成功率。

AI中文摘要:

视觉问答(VQA)与视觉语言模型(VLM)的结合日益广泛,尤其是在隐私敏感和带宽受限的场景中。联邦学习(FL)、分割学习(SL)和U形分割学习(USL)将原始数据保留在本地,但跨模型分区传输所有视觉令牌仍然成本高昂,并可能泄露隐私信息。我们提出QPriv-VL,一种面向FL、SL和USL的基于问题引导的隐私感知令牌剪枝框架,该框架在传输前根据任务效用和隐私敏感性对视觉令牌进行剪枝。其核心组件是一个轻量级动态阈值预测器(DTP),该预测器在一次前向传播中联合估计样本特定的剪枝比率和令牌级保留掩码。DTP将问题相关性(通过视觉补丁与池化问题嵌入之间的跨模态相似度计算)与从冻结的DINOv2特征中导出的敏感性信号相结合。这使得模型能够抑制潜在敏感区域,同时保留对回答问题有用的补丁,而无需敏感性标签。我们在GQA、OK-VQA、VQAv2、SLAKE、VQA-RAD和PathVQA上评估了QPriv-VL,并针对四类隐私攻击:FSHA、FORA、iDLG和属性推断成员推断攻击。DTP在匹配或超越固定比率剪枝的同时,显著减少了传输的令牌数量。在VQA-RAD上,它将成员推断攻击成功率从0.99降至0.76-0.79,相对于固定比率剪枝降低了FSHA和FORA重建PSNR,并使用约40%的原始视觉令牌预算保持了具有竞争力的VQA准确率。敏感性排除比率为1.20±0.18,表明优先移除了隐私敏感补丁,而可解释性分析显示,保留适应于问题语义而非通用视觉显著性。

英文摘要:

Visual Question Answering (VQA) with Vision-Language Models (VLMs) is increasingly used in privacy-sensitive and bandwidth-constrained settings. Federated Learning (FL), Split Learning (SL), and U-Shaped Split Learning (USL) keep raw data local, but transmitting all visual tokens across a model partition remains costly and can expose private information. We propose QPriv-VL, a question-guided, privacy-aware token-pruning framework for FL, SL, and USL that prunes visual tokens before transmission based on task utility and privacy sensitivity. Its core component is a lightweight Dynamic Threshold Predictor (DTP) that jointly estimates a sample-specific pruning ratio and a token-level retention mask in one forward pass. DTP combines question relevance, computed from cross-modal similarity between visual patches and the pooled question embedding, with a sensitivity signal derived from frozen DINOv2 features. This allows the model to suppress potentially sensitive regions while preserving patches useful for answering the question, without requiring sensitivity labels. We evaluate QPriv-VL on GQA, OK-VQA, VQAv2, SLAKE, VQA-RAD, and PathVQA against four privacy attack families: FSHA, FORA, iDLG, and attribute-inference membership inference attacks. DTP matches or outperforms fixed-ratio pruning while using substantially fewer transmitted tokens. On VQA-RAD, it reduces membership-inference attack success from 0.99 to 0.76-0.79, lowers FSHA and FORA reconstruction PSNR relative to fixed-ratio pruning, and preserves competitive VQA accuracy using about 40% of the original visual-token budget. A sensitivity exclusion ratio of 1.20 +/- 0.18 indicates preferential removal of privacy-sensitive patches, while explainability analysis shows that retention adapts to question semantics rather than generic visual saliency.

↑