GPEC:用于心脏超声视频字幕生成的高效预大语言模型高斯过程嵌入校正
GPEC: Efficient Pre-LLM Gaussian Process Embedding Correction for Cardiac Video Caption Generation
浏览论文内容
中文总结 AI 辅助
针对多模态大模型在心脏超声字幕生成中性能下降的问题,提出GPEC,一种插入视觉投影层与语言模型之间的高效预大语言模型残差校正方法,利用稀疏变分高斯过程改进视觉表示,提升字幕相似性与内容对齐,且推理开销极低。
中文摘要 AI 辅助
多模态大语言模型(MLLMs)在视频理解和字幕生成方面展现出强大的潜力,但在超声心动图等专业医学影像领域,其性能可能会下降。本工作引入了高斯过程嵌入校正(GPEC),一种模块化且计算高效的预大语言模型误差校正方法,用于改进VideoChat2在心脏超声字幕生成中所使用的视觉表示。GPEC插入在视觉投影层和语言模型之间,学习一个残差校正,将投影后的视觉表示移向一个由注释引导的目标。该目标通过将结构化视频注释转换为定性属性、生成固定格式的参考字幕,并将其映射到语言模型嵌入空间来构建。校正使用具有诱导点的稀疏变分高斯过程、自然参数变分更新和块状线性核进行建模,而原始VideoChat2组件保持冻结。该方法通过表示级、字幕级、内容导向和执行时间指标进行评估,在相同的输入和参考条件下比较原始VideoChat2与VideoChat2+GPEC。结果显示,应用所提出的校正后,字幕相似性和内容对齐性有所提高。此外,在评估设置中,GPEC每个视频增加的推理时间开销小于0.05秒。这些发现表明,GPEC能够以最小的计算成本改善专业医学视频领域的字幕生成,而无需对预训练的多模态骨干网络进行端到端微调。
英文摘要
Multimodal large language models (MLLMs) have shown strong potential for video understanding and caption generation, but their performance may decline in specialized medical imaging domains such as echocardiography. This work introduces Gaussian Process Embedding Correction (GPEC), a modular and computationally efficient pre-LLM error-correction method that improves the visual representations used by VideoChat2 for cardiac ultrasound caption generation. GPEC is inserted between the visual projection layer and the language model and learns a residual correction that moves the projected visual representation toward an annotation-guided target. The target is constructed by converting structured video annotations into qualitative attributes, generating a fixed-format reference caption, and mapping it into the language-model embedding space. The correction is modeled using a sparse variational Gaussian Process with inducing points, natural-parameter variational updates, and a block-wise linear kernel, while the original VideoChat2 components remain frozen.The method is evaluated using representation-level, caption-level, content-oriented, and execution-time metrics by comparing the original VideoChat2 with VideoChat2 + GPEC under identical input and reference conditions. Results show improved caption similarity and content alignment after applying the proposed correction. Furthermore, GPEC adds less than 0.05 s of inference-time overhead per video in the evaluated setting. These findings indicate that GPEC can improve caption generation in specialized medical video domains with minimal computational cost, without requiring end-to-end fine-tuning of the pretrained multimodal backbone.
发表机构
- K.N. Toosi University of Technology(K.N.图西理工大学)
机构由 AI 辅助整理,请以论文原文为准。