发表机构
The Chinese University of Hong Kong(香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出SemABR,利用多模态大语言模型定义视频语义保真度指标,并嵌入5G资源分配框架,在比特率受限时减少语义损失,优于传统方法。
AI 中文摘要
传统的视频指标如PSNR、SSIM和VMAF衡量视觉失真或感知质量,但它们并未直接捕捉语义保留:即压缩是否保留了视频中的对象、动作和时间叙事。现有的基于体验质量(QoE)的比特率选择和资源分配方法主要旨在最小化重缓冲和比特率切换,同时最大化感知视频质量,而未明确考虑语义保留。为解决这一空白,我们引入了视频语义保真度(SF),该指标量化压缩视频在多大程度上保留了其源视频的语义内容。一个离线的多模态大语言模型(MLLM)生成参考视频和压缩视频的结构化描述,另一个仅文本的大语言模型(LLM)评估它们的语义对应性。由此产生的依赖于内容的SF-比特率配置文件被缓存,并由在线比特率选择器查询,而无需在运行时调用MLLM。在三个主观QoE基准上的评估显示,SF与平均意见分数(MOS)之间存在一致的正相关。一项独立的人类语义评分研究评估了语义保留,并表明SF与人类判断的相关性强于传统视频指标。然后,我们将这些配置文件嵌入到基站处的5G MEC辅助视频点播(VoD)资源分配框架中。当无线资源无法为所有用户支持高比特率水平时,该框架使用SF-比特率配置文件跨用户联合选择比特率水平,并减少因所需比特率降低而导致的语义损失。使用5G NR模块的NS-3仿真表明,所提出的框架在平均和最差用户SF方面均优于评估的基线,并且随着每个用户可用的无线资源减少,其优势不断扩大。
英文摘要
Conventional video metrics such as PSNR, SSIM, and VMAF measure visual distortion or perceptual quality, but they do not directly capture semantic preservation: whether compression retains a video's objects, actions, and temporal narrative. Existing Quality-of-Experience (QoE)-driven bitrate-selection and resource-allocation methods primarily aim to minimize rebuffering and bitrate switching while maximizing perceptual video quality, without explicitly considering semantic preservation. To address this gap, we introduce video semantic fidelity (SF), a metric that quantifies how well a compressed video preserves the semantic content of its source. An offline multimodal large language model (MLLM) generates structured descriptions of the reference and compressed versions of the video, and a separate text-only large language model (LLM) evaluates their semantic correspondence. The resulting content-dependent SF--bitrate profiles are cached and queried by the online bitrate selector without invoking MLLMs at runtime. Evaluations on three subjective QoE benchmarks show a consistent positive association between SF and mean opinion scores (MOS). A separate human semantic-rating study evaluates semantic preservation and shows that SF correlates more strongly with human judgments than conventional video metrics. We then embed these profiles into a 5G MEC-assisted video-on-demand (VoD) resource-allocation framework at the base station. When wireless resources cannot support high bitrate levels for all users, the framework uses the SF--bitrate profiles to select bitrate levels jointly across users and reduce the semantic loss caused by the required bitrate reductions. NS-3 simulations with the 5G NR module show that the proposed framework achieves higher average and worst-user SF than the evaluated baselines, with a widening advantage as the wireless resources available to each user decrease.