arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型(LLM)能否设计视频编码工具?以平面模式为例

Can LLMs Design Video Coding Tools? A Case Study on Planar Mode

Yingwen Zhang, Meng Wang, Liqiang He, Shiqi Wang

arXiv 2609.01535首次发表:更新:

发表机构

City University of Hong Kong; Lingnan University(香港城市大学; 岭南大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过案例探究LLM能否设计视频编码工具,以平面模式为对象,借助生成-评估循环优化,在VVenC和ECM上实现码率节省,凸显其潜力与挑战。

AI 中文摘要

本文探索大语言模型(LLM)是否能设计视频编码工具,该任务极具挑战性,因为工具修改存在复杂的算法耦合。我们针对视频编码标准中历史悠久的帧内预测工具——平面(Planar)模式,开展实证案例研究。实验采用生成-评估循环:LLM生成新的平面预测器,编码器进行试验以评估其编码性能,LLM则根据评估反馈重新生成优化后的实现。我们首先在弗劳恩霍夫通用视频编码器(VVenC)的快速预设设置下,直接替换默认平面模式,实验结果表明,在该轻量工具集上,LLM生成的模式优于传统平面模式,在标准基准上实现0.18%的码率节省,同时带来0.4%的复杂度开销。我们进一步将评估扩展至增强压缩模型(ECM),利用新引入的定向平面模式,研究两种集成策略:直接替换它们,或将LLM生成的预测器作为带有新语法元素的额外预测模式引入。实证结果显示,在受限低分辨率设置下,两种策略均可产生编码增益。总体而言,本研究提供了初步证据与实践见解,凸显了基于LLM的编码工具设计的潜力与未解决的挑战。

英文摘要

This paper explores whether large language models (LLMs) can design video coding tools, a highly challenging task due to the intricate algorithmic coupling of tool modifications. In particular, we present an empirical case study on the Planar mode, a long-standing intra prediction tool in video coding standards. Our experiments operate within a generation-and-evaluation loop, with the LLM generating new Planar predictors, encoder trials evaluating their coding performance, and the LLM re-generating refined implementations based on the evaluation feedback. We first examine directly replacing the default Planar mode in the Fraunhofer Versatile Video Encoder (VVenC) under its faster preset. Experimental results demonstrate that the LLM-generated mode can outperform the conventional Planar mode on this lightweight toolset, achieving 0.18% bitrate savings with 0.4% complexity overhead on the standard benchmark. We further extend our evaluation to the Enhanced Compression Model (ECM). Leveraging newly introduced directional Planar modes, we investigate two integration strategies: directly replacing them, and introducing the LLM-generated predictor as an additional prediction mode with new syntax elements. The empirical results suggest that both strategies can yield coding gains under a constrained low-resolution setting. Overall, this study offers preliminary evidence and practical insights, highlighting both the potential and open challenges of LLM-based coding tool design.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑