发表机构
Brandeis University(布兰迪斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对因带宽限制导致的低质量视频通话问题,提出FSFVE框架,利用仅10帧面部图像快速训练模型,可在通话前或期间训练,能在低计算设备上实时运行,显著提升压缩面部视频质量。
AI 中文摘要
视频通话如今已成为一种流行的通信方式,许多公司提供免费服务。然而,由于带宽限制,全球仍有数百万人经历低质量视频通话,即便大多数人具备所需硬件。本文提出一种用于增强高度压缩视频通话的新颖框架。研究表明,仅用10帧面部图像,就能在100秒内快速训练模型来增强该视频通话实例。模型可在通话前或通话期间训练,通过生成更高质量视频来增强通话其余部分。视频会议应用无需修改,系统作为顶层可快速训练,然后让视频会议应用照常运行,在图像显示前进行拦截和改进。该模型设计为能在典型笔记本电脑CPU等低计算设备上实时运行。实验表明,该模型在量化和感知方面都显著提高了压缩面部视频的质量。代码可在指定网址获取。
英文摘要
Videocalling has become a popular form of communication in the world today, with many companies providing free services for it. However, there are still millions of people around the world that experience poor quality videocalls due to limitations in bandwidth. This despite, most people having the required hardware. In this paper we present a novel framework for enhancing highly compressed videocalls. We show, that with as little as 10 frames of the face, we can rapidly (in under 100 seconds) train a model to enhance that instance of the videocall. The model can be trained either prior to or during the call, enhancing the rest of the call by producing better quality video. The video conferencing application need not be modified - it can be off the shelf with our system as a layer on top that trains quickly then simply lets the video conferencing application (e.g. Zoom) run as usual, where our system intercepts and improves images before they are displayed. The model is designed to run in realtime on low-compute devices such as a typical laptop CPU. Experimentally, we show that the model significantly improves quality of compressed face video both quantitatively as well as perceptually. Code can be found at https://github.com/varun-jois/FSFVE.
DOI:10.1109/DCC66757.2026.00021