发表机构
SAMOVAR Laboratory, Télécom SudParis, Institut Polytechnique de Paris; Moments Lab(萨莫瓦尔实验室,巴黎电信学院,巴黎综合理工学院; Moments Lab公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本综述系统梳理视频与视听大语言模型的推理效率优化机制,按流水线阶段组织帧采样、模态编码、标记缩减及预填充解码等方法,汇总共享条件下的精度-成本对比,指出视听效率与标准化评估的空白。
AI 中文摘要
视频理解已迅速向视频大语言模型(VideoLLMs)方向发展:这类系统将视频表示与预训练的大语言模型相结合,并根据文本提示进行条件生成。它们在字幕生成、问答、检索和时间定位方面的强大性能,是以随帧数和上下文长度增长的计算和内存成本为代价的,这限制了其在实时、移动和资源受限场景中的部署。本综述涵盖了视觉和视听VideoLLMs的推理效率机制,这些机制报告了在参数数量、每输入FLOPs、延迟、内存或视觉和音频标记数量方面的具体减少。我们分析了帧采样、模态编码、连接器级标记减少以及LLM预填充和解码等环节的瓶颈。我们按方法作用的流水线阶段对其进行组织,涵盖了自2022年底以来开发的VideoLLMs,以及仍作为当前流水线组成部分的早期帧采样和视觉编码器机制。我们在共享宿主模型和输入协议下汇集了文献报告的准确率-成本比较(只要可得),将其与异构跨论文证据区分开来,并识别了视听效率和标准化评估方面的空白。我们在以下网址维护了一个仓库:此https URL。
英文摘要
Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video representations with pretrained large language models and condition generation on a textual prompt. Their strong performance on captioning, question answering, retrieval and temporal grounding comes at a computation and memory cost that grows with frame count and context length, limiting deployment in real-time, mobile and resource-constrained settings. This survey covers inference-efficiency mechanisms for visual and audiovisual VideoLLMs that report concrete reductions in parameter count, FLOPs per input, latency, memory, or visual and audio token count. We analyze bottlenecks across frame sampling, modality encoding, connector-level token reduction, and LLM prefilling and decoding. We organize methods by the pipeline stage at which they act, covering VideoLLMs developed since late 2022 together with earlier frame-sampling and vision-encoder mechanisms that remain components of current pipelines. We assemble literature-reported accuracy--cost comparisons under shared host models and input protocols wherever available, distinguish them from heterogeneous cross-paper evidence, and identify gaps in audiovisual efficiency and standardized evaluation. We maintain a repository at https://github.com/momentslab/awesome-efficient-videollm.
CommentsSupplementary material at https://www.killian-steunou.com/videollm-survey/static/pdfs/videollm_survey_supplementary.pdf