高保真视频质量评估与VQA特定显著性
High-Fidelity Video Quality Assessment with VQA-Specific Saliency
- The University of Texas at Austin(德克萨斯大学奥斯汀分校)
- University of Colorado Boulder(科罗拉多大学博尔德分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出HFVQA框架,利用固定大小时空补丁与VQA特定显著性,在保留高保真线索的同时仅处理12%补丁,实现无参考视频质量评估的最先进性能。
AI中文摘要:
无参考视频质量评估(NR VQA)近年来随着深度学习的发展取得了显著的进展。然而,视频数据本质上规模庞大,使用深度模型处理这些数据会产生高昂的计算成本。这一挑战在VQA中尤为突出,因为保留原始分辨率线索和密集的时间信息对于准确性至关重要。现有的效率驱动预处理策略,如分片处理,虽然减少了计算量,但改变了输入数据分布,限制了预训练视频基础模型(ViFMs)的有效复用。为解决这些问题,我们提出了高保真视频质量评估(HFVQA),这是一个基于固定大小时空(ST)补丁的框架,与预训练的ViFMs完全兼容。HFVQA在多个尺度上采样ST补丁,包括原始分辨率,并进行最小的时间子采样以保留低级质量线索和语义上下文。为限制计算量,HFVQA引入了一个轻量级辅助网络,与ViFM编码器端到端训练,以学习VQA特定显著性。该显著性直接从质量监督中蒸馏得到,捕获任务特定的重要性模式,反映出视频质量感知主要由一小部分时空区域主导。通过将高保真时空线索与学习到的任务特定显著性相结合,HFVQA在标准NR VQA基准上实现了最先进的性能,同时仅处理12%的候选ST补丁,使得基于ViFM的高保真VQA在计算上变得可行。
英文摘要:
No-reference video quality assessment (NR VQA) has recently seen promising progress with deep learning. However, video data is inherently large, and processing them with deep models incurs high computational cost. This challenge is particularly acute in VQA, where preserving original-resolution cues and dense temporal information is critical for accuracy. Existing efficiency-driven preprocessing strategies, such as fragmenting, reduce computation but alter the input data distribution, limiting effective reuse of pretrained video foundation models (ViFMs). To address these challenges, we propose \textbf{H}igh-\textbf{F}idelity \textbf{V}ideo \textbf{Q}uality \textbf{A}ssessment (\textbf{HFVQA}), a framework built on fixed-size spatio-temporal (ST) patches that is fully compatible with pretrained ViFMs. HFVQA samples ST patches across multiple scales, including the original resolution, with minimal temporal subsampling to preserve low-level quality cues and semantic context. To limit computation, HFVQA introduces a lightweight auxiliary network trained end-to-end with the ViFM encoder to learn \textit{VQA-specific saliency}. Distilled directly from quality supervision, this saliency captures task-specific importance patterns, reflecting that video quality perception is dominated by a small subset of spatio-temporal regions. By combining high-fidelity spatio-temporal cues with learned, task-specific saliency, HFVQA achieves SOTA performance on standard NR VQA benchmarks while processing as little as 12\% of candidate ST patches, making high-fidelity ViFM-based VQA computationally tractable.