AI 中文总结
该研究针对协作式LVLM推理的隐私风险,提出RASR多模态重建攻击,可从深层LVLM隐藏状态恢复隐私信息,图像重建MSE降约50%、文本恢复token准确率达99%。
AI 中文摘要
协作式推理通过在边缘设备与云之间划分计算任务来部署大型视觉语言模型(LVLM)。尽管保留原始输入被认为能保障隐私,但传输中间隐藏状态会暴露关键攻击面。然而,鉴于视觉内容已被投影到语言嵌入空间,深层LVLM隐藏状态是否仍保留可恢复的私人信息尚不明确。为解决这一问题,我们从理论上分析了LVLM隐藏状态的可恢复性,发现在正则性假设和正语义-干扰边际下,与隐私相关的视觉语义仍可识别且稳定可恢复。基于此分析,我们提出RASR,一种新颖的粗到细多模态重建攻击方法。RASR通过遵循各自正向处理管道的反向路径,利用特定模态的逆路径获取初始图像和文本重建结果,再借助隐藏状态一致性对两种重建结果进行优化。在Qwen3-VL-8B-Instruct和LLaVA-1.5-7B模型上,针对五个数据集的评估显示,与最强基线相比,RASR将图像重建MSE降低了约50%,同时实现了高达99%的文本恢复token准确率。这些结果表明,即使从深层LVLM隐藏状态中也能恢复出隐私敏感的视觉和文本信息,揭示了协作式推理的隐私风险。
英文摘要
Collaborative inference deploys Large Vision-Language Models (LVLMs) by partitioning computation between edge devices and the cloud. While withholding raw inputs supposedly ensures privacy, transmitting intermediate hidden states exposes a critical attack surface. However, it remains unclear whether deep-layer LVLM hidden states retain recoverable private information, given that visual content has been projected into the language embedding space. To address this concern, we theoretically analyze LVLM hidden-state recoverability and show that, under regularity assumptions and a positive semantic--nuisance margin, privacy-relevant visual semantics remain identifiable and stably recoverable. Motivated by this analysis, we propose RASR, a novel coarse-to-fine multimodal reconstruction attack. RASR obtains initial image and text reconstructions through modality-specific inverse paths that follow their respective forward processing pipelines in reverse, and then uses hidden-state consistency to refine both reconstructions. Evaluations on Qwen3-VL-8B-Instruct and LLaVA-1.5-7B across five datasets demonstrate that RASR reduces image reconstruction MSE by \(\sim\)50\% compared to the strongest baselines, while achieving up to 99\% token accuracy for text recovery. These results show that privacy-sensitive visual and textual information can be recovered even from deep-layer LVLM hidden states, exposing the privacy risks of collaborative inference.