arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28691cs.CVcs.AI

防御可穿戴视觉语言模型(VLM)抵御私有属性推断

Defending Wearable VLMs Against Private Attribute Inference

Zhimin Li, Pan Wang, Jingxian Chen, Yuantao Tang, Anthony Chen, Qian Lou, Jingtong Hu

首次发表
浏览论文内容

中文总结 AI 辅助

针对可穿戴VLM的视觉令牌可能泄露私有属性的问题,提出TGAP方法,可大幅降低隐私推断准确率,同时维持较高效用,为可穿戴多模态AI的隐私保护提供可行路径。

中文摘要 AI 辅助

可穿戴VLM流程可从自我中心视觉捕获中提供连续多模态辅助:用户对周围场景提出任务驱动的问题,系统使用紧凑视觉令牌支持语言推理。本研究的核心挑战在于,辅助所需的自我中心证据可能泄露佩戴者或附近旁观者的私有属性。我们将此视为拆分VLM推理的隐私-效用联合问题,其中视觉编码在可信设备边界内进行,但中间视觉令牌可能被传输到下游推理组件,这暴露了未被充分研究的泄露面:即使最终文本响应是良性的,外部攻击者或不可信下游组件仍可从传输的视觉令牌中恢复私有属性。为评估这种权衡,我们构建了包含3221条图像-问题记录的配对隐私-效用基准,每条记录对应一个效用问题及涵盖位置、收入、性别和兴趣的隐私标签。我们进一步提出Token-Guided Attribute Privacy(TGAP),一种预大型语言模型(LLM)的令牌解缠器,在视觉令牌离开可信边界前对其进行残差变换学习。TGAP结合效用保留、身份正则化、语义隐私抑制和图像驱动表示抑制,避免了粗糙硬掩码或注意力掩码导致的效用损失。在源模型评估所用基准上,TGAP将隐私准确率从56.7%降至7.4%,绝对下降49.3%,同时保持74.4%的宽松效用。这些结果表明,保护紧凑令牌接口是实现隐私保护型可穿戴多模态AI的可行路径。

英文摘要

Wearable VLM pipelines promise continuous multimodal assistance from egocentric visual capture: a user asks a task-driven question about the surrounding scene, and the system uses compact visual tokens to support language reasoning. The challenge motivating this work is that the same egocentric evidence needed for useful assistance can also reveal private attributes about the wearer or nearby bystanders. We investigate this as a joint privacy-utility problem for split VLM inference, where visual encoding occurs within a trusted device boundary but intermediate visual tokens may be transmitted to downstream reasoning components. This exposes an understudied leakage surface: even when final textual responses are benign, external attackers or untrusted downstream components can recover private attributes from transmitted visual tokens. To evaluate this tension, we construct a paired privacy-utility benchmark with 3,221 image-question records, each paired with a utility question and privacy labels covering location, income, sex, and interests. We further propose Token-Guided Attribute Privacy (TGAP), a pre-LLM token disentangler that learns a residual transformation of visual tokens before they leave the trusted boundary. TGAP combines utility preservation, identity regularization, semantic privacy suppression, and image-driven representation suppression, avoiding the utility loss caused by coarse hard or attention masking. On the benchmark used for source-model evaluation, TGAP reduces privacy accuracy from 56.7\% to 7.4\%, a 49.3\% absolute drop, while maintaining relaxed utility at 74.4\%. These results suggest that securing the compact token interface is a practical path toward privacy-preserving wearable multimodal AI.

发表机构

  • Swanson School of Engineering, University of Pittsburgh(匹兹堡大学斯旺森工程学院)
  • North Allegheny Senior High(北阿勒格尼高级中学)
  • College of Engineering, Michigan State University(密歇根州立大学工程学院)
  • University of Central Florida(中佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

↑