arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09139cs.CV

CodecArena:基于视觉强化学习的编解码器质量评估

CodecArena: Codec Quality Assessment via Visual Reinforcement Learning

Jiaye Fu, Weiqi Li, Qiankun Gao, Yanchen Zhao, Xiandong Meng, Jian Zhang, Siwei Ma, Jiaqi Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对视频编码主流指标无法准确评估内容保真度的问题,提出首个视频编码质量评估视觉-语言框架CodecArena,采用视觉强化学习方案优化,构建相关数据集与基准,在源不相交内容上实现优于现有方法的人类判断一致性。

中文摘要 AI 辅助

视频编码正朝着低比特率和超比特率领域发展,这得益于端到端编解码器(将手工设计的流水线替换为联合优化的神经网络)和生成式编解码器(利用视频生成模型的先验知识)的进步。然而,主流指标LPIPS和DISTS仅测量特征和纹理相似性,而非内容保真度:即使重构出的图像产生错误人脸或将文本模糊为看似合理的笔画,也能获得高分,而人类会立即拒绝此类结果。为解决这一问题,我们提出CodecArena,这是首个用于视频编码质量评估的视觉-语言框架,将编解码器评估转化为参考图像与其重构结果之间基于源的比较推理。我们采用Facet-GRPO优化CodecArena,这是一种视觉强化学习方案,可在成对编解码器偏好对齐的同时,将判决依据锚定在五个保真度维度:身份、物体、文本、纹理和时间一致性。其锚定维度的奖励使用自动推导的维度方向作为弱锚点,而非人类的各维度标签,以防止任何单一子分数主导整体偏好,并产生可解释的细粒度质量判断。为支持这一未充分探索领域的训练与评估,我们构建了两个互补资源:CodecArena-1K,一个包含1500个比较组的全自动偏好数据集,由传统、神经和生成式编解码器的重构结果构建,融合了视觉-语言和客观监督;CodecArena-Bench,一个具有源不相交视频的人类排序基准,用于公平的域外评估。大量实验表明,CodecArena在源不相交内容上对人类判断的一致性达到了最先进水平,覆盖不同编解码器和比特率,优于感知指标和先前的视觉-语言评估器。

英文摘要

Video coding is advancing into the low and ultra-low bitrate regime, driven by end-to-end codecs that replace the hand-crafted pipeline with jointly optimized neural networks and generative codecs that exploit the priors of video generation models. Yet the dominant metrics, LPIPS and DISTS, measure feature and texture similarity rather than content fidelity: a reconstruction that hallucinates a wrong face or blurs text into convincing strokes can still score well, even when a human rejects it instantly. To address this, we propose CodecArena, the first vision-language framework for video coding quality assessment, casting codec evaluation as source-conditioned comparative reasoning between a reference and its reconstructions. We optimize CodecArena with Facet-GRPO, a visual reinforcement learning scheme that aligns pairwise codec preferences while grounding the verdict in five fidelity facets: identity, objects, text, texture, and temporal consistency. Its facet-anchored reward uses automatically derived facet directions as weak anchors, rather than human per-facet labels, to prevent any single sub-score from dominating the holistic preference and to yield interpretable fine-grained quality judgments. To support training and evaluation in this underexplored regime, we construct two complementary resources: CodecArena-1K, a fully automatic preference dataset of 1,500 comparison groups built from traditional, neural, and generative codec reconstructions with fused vision-language and objective supervision; and CodecArena-Bench, a human-ranked benchmark with source-disjoint videos for fair out-of-domain evaluation. Extensive experiments demonstrate that CodecArena achieves state-of-the-art agreement with human judgments on source-disjoint content across diverse codecs and bitrates, surpassing perceptual metrics and prior vision-language evaluators.

补充信息

↑