arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20212cs.CV

校正镜头:一种基于物理的视频眼镜去除方法

Unwarping the Lens: A Physics-Grounded Approach to Video Glasses Removal

Radim Spetlik, David Futschik, Radek Danecek, Feitong Tan, Ziqian Bai, Rohit Pandey, Yinda Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

提出基于物理的迁移框架,将 Nano Banana 等模型知识迁移至 JFSnet,在 FFHQ 等数据集上实现高保真视频眼镜去除,性能优于扩散与 GAN 基线。

中文摘要 AI 辅助

从视频中高保真地去除眼镜是面部属性编辑中的一项重大挑战,因为底层面部几何结构常被复杂的折射畸变和视角相关的镜面反射所掩盖。尽管大规模生成先验通过静态图像修复在眼镜去除方面展现出潜力,但它们往往缺乏维持身份、表情和姿态所需的结构约束,导致静态图像和动态序列中出现明显的“身份漂移”。本文提出一种新型迁移框架以解决生成先验的随机性问题:该流程首先从商用级生成模型 Nano Banana、Gemini 3 Pro Image 中提取高保真合成人脸图像,通过三阶段结构过滤过程对其进行正则化以保留身份、表情和姿态,最终在训练期间应用基于物理的镜头光学模拟以提供多样化配对数据。此过程将 Nano Banana 的照片级真实感多视图知识迁移至专用恢复架构 JFSnet(Joint Feature-Spatial network),JFSnet 集成基于 DINOv2 的语义特征与用于空间重建的卷积解码器,利用平移等变性约束提升时间一致性和高频细节保留能力。在精选的 Flickr-Faces-HQ(FFHQ)子集(12163 张图像)上的评估显示,该方法实现了高保真度和结构准确性,同时保持 27.68 FPS 的推理速度;在 CelebV-Text 视频序列的感知研究中,其结果在眼部一致性、时间稳定性和整体恢复质量上均优于基于扩散和 GAN 的基线方法。

英文摘要

High-fidelity removal of eyeglasses from video is a major challenge in facial attribute editing, as the underlying facial geometry is often obscured by complex refractive distortions and view-dependent specular reflections. While large-scale generative priors have shown promise in eye-glasses removal via static image inpainting, they often lack the structural constraints necessary to maintain identity, expression, and pose, leading to visible "identity drift" in both static images and dynamic sequences. In this paper, we propose a novel transfer framework that addresses the stochastic nature of generative priors. Our pipeline first extracts high-fidelity synthetic face images from a commercial-grade generative model (Nano Banana, Gemini 3 Pro Image), regularizes them via a three-stage structural filtering process to preserve identity, expression, and pose, and finally applies physically-based simulation of lens optics during training to provide diverse, paired data. This process transfers Nano Banana's photo-realistic, multi-view knowledge into a specialized restoration architecture, JFSnet (Joint Feature-Spatial network). JFSnet integrates DINOv2-based semantic features with a convolutional decoder for spatial reconstruction, leveraging translation equivariance constraints to improve temporal consistency and high-frequency detail preservation. Evaluations on the curated Flickr-Faces-HQ (FFHQ) subset (12,163 images) show that our approach achieves high fidelity and structural accuracy, while maintaining inference speed of 27.68 FPS. In perceptual studies on CelebV-Text video sequences, our results are consistently preferred over diffusion and GAN-based baselines for ocular consistency, temporal stability, and overall restoration quality.

发表机构

  • Czech Technical University in Prague(布拉格捷克技术大学)
  • Faculty of Electrical Engineering(电气工程学院)
  • Google(谷歌公司)

机构由 AI 辅助整理,请以论文原文为准。

↑