arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11645cs.CV

隐形斗篷:实时隐私保护的体视频流

Cloak of Invisibility: Real-Time Privacy-Preserving Volumetric Video Streaming

  • University of California Los Angeles(加利福尼亚大学洛杉矶分校)
  • Nokia Bell Labs(诺基亚贝尔实验室)

机构由 AI 辅助整理,请以论文原文为准。

Hossein Khalili, Philip Do, Alexander Vilesov, Achuta Kadambi, Kittipat Apicharttrisorn, Nader Sehatbakhsh

AI总结:

针对体视频流的隐私挑战,提出实时源端隐私系统InViStream,结合目标检测与深度感知掩码等技术,在多视角场景中实现高效隐私保护与实时重建,取得优异性能指标。

AI中文摘要:

体视频流将隐私问题转化为三维多视角问题。与普通视频中可逐帧编辑敏感内容不同,RGB-D体视频管线通过多台相机捕捉人物、房间和个人物品,并将其融合为共享的三维表示。某一视角中遗漏的私有对象,或在融合前仅被部分移除的对象,会在重建场景中重新出现。这为三维远程呈现、教育、娱乐和沉浸式应用带来了隐私挑战:需在原始视觉和几何数据离开相机侧前移除私有内容,同时场景的公共部分需保持对实时重建的可用性。现有体视频流系统主要优化重建、数据传输和延迟,而隐私保护视觉方法针对单相机单帧图像设计,无法直接处理校准后的多视角RGB-D融合。我们提出InViStream,一种适用于该场景的实时“源端隐私”系统。InViStream解决体捕捉中的三个挑战:私有对象在不同视角中呈现可能不同;仅RGB掩码会在深度上留下几何隐私泄露;同一类别的公共/私有实例需在云端融合前一致分离。为应对这些挑战,InViStream将目标检测与深度感知掩码结合,在校准视角间传播公共/私有决策,仅融合经净化的点云。我们在合成和真实RGB-D场景(包括办公室、会议室、客厅,以及含多个公共和私有人员与对象的场景)上评估InViStream。InViStream在合成数据上实现Dice/召回率为0.799/0.891,在真实数据上实现Dice/召回率为0.792/0.908,合成数据的结构相似性指数(SSIM)高于0.98,且实时流传输帧率超过30 FPS。

英文摘要:

Volumetric video streaming turns privacy into a 3D, multi-view problem. Unlike ordinary video, where sensitive content can often be redacted frame by frame, RGB-D volumetric pipelines capture people, rooms, and personal objects from multiple cameras and fuse them into a shared 3D representation. A private object missed in one view, or only partially removed before fusion, can therefore reappear in the reconstructed scene. This creates a privacy challenge for 3D telepresence, education, entertainment, and immersive applications: private content should be removed before raw visual and geometric data leave the camera side, while the public part of the scene should remain useful for real-time reconstruction. Existing volumetric streaming systems mainly optimize reconstruction, data movement, and latency, while privacy-preserving vision methods are designed for single-camera, single-frame images and do not directly address calibrated multi-view RGB-D fusion. We present InViStream, a real-time "privacy-from-source" system designed for this setting. InViStream addresses three challenges in volumetric capture: private objects may appear differently across views, RGB masking alone can leave geometric privacy leakage in depth, and public/private instances of the same class must be separated consistently before cloud-side fusion. To address these challenges, InViStream combines object detection with depth-aware masking, propagates public/private decisions across calibrated views, and fuses only sanitized point clouds. We evaluate InViStream on synthetic and real RGB-D scenes, including offices, conference rooms, living rooms, and settings with multiple public and private people and objects. InViStream achieves synthetic Dice/Recall of 0.799/0.891 and real Dice/Recall of 0.792/0.908, with synthetic SSIM above 0.98 and real-time streaming above 30 FPS.

↑