arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

为视觉几何接地Transformer生成多视角对抗样本

Generating Multi-view Adversarial Examples for Visual Geometry Grounded Transformer

Qi Song, Ziyuan Luo, Haoliang Han, Renjie Wan

arXiv 2608.20748首次发表:更新:

发表机构

Hong Kong Baptist University(香港浸会大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对3D基础模型VGGT的安全漏洞,提出MVAP-G多视角对抗扰动生成器,通过跨视角对抗对齐机制生成一致扰动,可无需迭代优化即可显著降低VGGT性能,为3D视觉系统鲁棒性研究提供了新方向。

AI 中文摘要

视觉几何接地Transformer(Visual Geometry Grounded Transformer,简称VGGT)可实现从多视角图像进行统一的前馈式3D重建。然而,部署这类高性能模型可能会暴露关键的安全漏洞。传统对抗扰动需要针对每个场景进行成本高昂的优化,而通用对抗扰动(Universal Adversarial Perturbations,简称UAPs)依赖单一静态模式,无法有效攻击VGGT。为解决这些局限,我们提出MVAP-G,一种多视角对抗扰动生成器,可在单次前馈过程中生成跨多个视角的不可察觉的一致扰动。为确保不同场景下的扰动一致性,我们设计了跨视角对抗对齐机制来处理多视角图像。实验表明,MVAP-G在推理阶段无需迭代优化即可显著降低VGGT的性能。本研究开创了针对3D基础模型的多视角对抗攻击,揭示了严重漏洞,强调了开发鲁棒3D视觉系统的迫切需求。代码可在该https URL获取。

英文摘要

The Visual Geometry Grounded Transformer (VGGT) enables unified feed-forward 3D reconstruction from multi-view images. However, deploying such a high-performance model may expose critical security vulnerabilities. Traditional adversarial perturbations require costly per-scene optimization, while Universal Adversarial Perturbations (UAPs) rely on a single static pattern and fail to effectively attack VGGT. To address these limitations, we propose \textbf{MVAP-G}, a multi-view adversarial perturbation generator that produces imperceptible consistent perturbations across multiple views in a single feed-forward pass. To ensure perturbation consistency across diverse scenes, we design a cross-view adversarial alignment mechanism to process multi-view images. Experiments demonstrate that MVAP-G significantly degrades VGGT performance without iterative optimization during inference. This work pioneers multi-view adversarial attacks on 3D foundation models, uncovering severe vulnerabilities and underscoring the urgent need for robust 3D vision systems. The code is available at https://github.com/qsong2001/mvap-g.

CommentsECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑