AI 中文总结
针对部分伪造视频检测难题,提出UVIF框架,通过统一编码器、伪标注及特征对齐策略联合建模视频与图像,检测性能优于现有最优方法且无额外计算开销。
AI 中文摘要
鉴于人脸操纵技术和深度生成模型的快速发展,人脸伪造检测对于保障人脸数据的安全性和完整性至关重要。现有的视频人脸伪造检测方法通常假设伪造视频中的所有帧均被操纵,然而检测仅包含部分篡改帧的部分伪造视频仍然具有挑战性。为解决该问题,我们提出了一种新框架UVIF,利用额外的带标注图像为视频中的部分伪造检测提供细粒度监督。UVIF采用统一编码器和多任务学习范式,联合建模人脸视频与图像以提升视频人脸伪造检测性能;其中统一编码器采用带时间融合模块的2D骨干网络。我们设计了针对视频帧的伪标注流程,以桥接视频帧与静态图像的表示;还引入了面向视频的特征对齐策略,以缩小视频与图像之间的分布差距。在基准数据集上的大量实验表明,我们的框架在检测部分伪造视频时优于现有最优方法,且未引入额外计算开销。我们的代码可在该https URL获取。
英文摘要
Face forgery detection is crucial for preserving the security and integrity of facial data given the rapid developments in face manipulation techniques and deep generative models. Existing methods for video face forgery detection typically assume that all frames in a forged video are manipulated, while detecting partially forged videos that contain only a subset of altered frames remains challenging. To address this issue, we propose a novel framework, UVIF, that utilizes additional annotated images to provide fine-grained supervision for detecting partial forgeries in videos. UVIF employs a unified encoder and a multi-task learning paradigm to jointly model facial videos and images for boosted video face forgery detection. A 2D backbone with temporal fusion modules is employed as the unified encoder. A pseudo labeling process is designed for video frames to bridge their representations with those of static images. A video-oriented feature alignment strategy is further introduced to reduce the distribution gap between videos and images. Extensive experiments on benchmark datasets demonstrate the effectiveness of our framework, which outperforms state-of-theart methods in detecting partially forged videos while introducing no additional computational overhead. Our code is available at https://github.com/haotianll/UVIF.