FANVIDv2:在复合退化下通过人脸和车牌识别评估视频超分辨率
FANVIDv2: Evaluating Video Super-Resolution by Face and Licence-Plate Recognition Under Compound Degradation
浏览论文内容
中文总结 AI 辅助
FANVIDv2基准通过人脸和车牌识别指标评估视频超分辨率,采用复合退化生成低分辨率视频,实验表明RCDM基线显著提升识别性能。
中文摘要 AI 辅助
视频超分辨率(VSR)通常通过在双三次下采样的片段上计算PSNR和SSIM来评判,尽管在监控场景中,其目的是使人脸和车牌变得“可识别”。我们提出了FANVIDv2,一个通过识别管道对其输出所能完成的任务来给VSR评分的基准。FANVIDv2提供了320×180分辨率的低分辨率(LR)片段,并带有48位公众人物(每人一张高分辨率画廊图像)和375个车牌片段(覆盖360个不同车牌字符串)的高分辨率(HR)参考。LR片段是通过随机复合退化(模糊、缩放抖动、传感器噪声、JPEG压缩、最终下采样)生成的,而非仅使用双三次下采样。两个指标对检测区域内的识别进行评分:FaceRecBox仅在人脸被定位并正确识别时给予奖励,而TextRecBox通过按定位质量加权的归一化编辑距离对车牌转录进行评分。使用一个2.3M参数的VSR基线(RCDM),FaceRecBox从0.6864提升到0.7222,匹配人脸的识别准确率从84.35%提升到86.93%,TextRecBox从0.3088提升到0.3667;一个残差图门控变体(RCDM-RMGF)在车牌上达到0.3801。我们详细描述了退化模型、基线架构和评分器,并发布了标注、元数据、下载和退化脚本以及评估代码。
英文摘要
Video super-resolution (VSR) is normally judged by PSNR and SSIM on clips that were downsampled bicubically, although in surveillance its purpose is to make faces and licence plates \emph{recognisable}. We present FANVIDv2, a benchmark that scores VSR by what a recognition pipeline can do with its output. FANVIDv2 provides $320\times180$ low-resolution (LR) clips with high-resolution (HR) references for 48 public figures (with one HR gallery image each) and 375 licence-plate clips covering 360 distinct plate strings. LR clips are generated with a randomised compound degradation (blur, resize jitter, sensor noise, JPEG compression, final downsampling) rather than bicubic downsampling alone. Two metrics score recognition \emph{inside} detections: FaceRecBox rewards a face only if it is localised and correctly identified, and TextRecBox scores plate transcriptions by normalised edit distance weighted by localisation quality. With a 2.3\,M-parameter VSR baseline (RCDM), FaceRecBox rises from 0.6864 to 0.7222, identity accuracy on matched faces from 84.35\% to 86.93\%, and TextRecBox from 0.3088 to 0.3667; a residual-map gated variant (RCDM-RMGF) reaches 0.3801 on plates. We describe the degradation model, the baseline architectures and the scorers in detail, and release annotations, metadata, download and degradation scripts and evaluation code.
发表机构
- Indian Institute of Technology Bombay(印度理工学院孟买分校)
机构由 AI 辅助整理,请以论文原文为准。