发表机构
Beijing University of Posts and Telecommunications; Peking University; Beijing Wuzi University(北京邮电大学; 北京大学; 北京物资学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对无人机视频理解任务,提出多模态大语言模型SkyVLaM,通过时间基感知器构建稀疏令牌,正则化稀疏基,自适应选择密集段,联合大语言模型处理,还构建SkyVid,实验证明其能有效分配视觉令牌预算,提升语言条件视频分割效果。
AI 中文摘要
多模态大语言模型(MLLMs)的最新进展显著提升了遥感(RS)多模态理解。语言条件分割对无人机(UAV)视频中的细粒度目标理解至关重要。然而,由于小的、视觉上模糊的目标以及动态空中视角的普遍存在,该任务仍具有挑战性。本文提出了SkyVLaM,一种用于无人机视频理解的多模态大语言模型。它通过时间基感知器直接从补丁级视频表示构建稀疏令牌,对稀疏基进行正则化以鼓励互补时间线索,并自适应选择时间连贯的密集段进行高分辨率检查。生成的稀疏和密集令牌由大语言模型联合处理以进行查询条件分割。还构建了SkyVid,包括SkyVid-VGCG和SkyVid-RVOS分别用于视频基础对话生成和引用视频对象分割。SkyVid包含101个视频、33.6K帧和153万个像素级对象实例。实验表明SkyVLaM在无人机场景中能更有效地分配视觉令牌预算并改善语言条件视频分割。
英文摘要
Recent advances in Multimodal Large Language Models (MLLMs) have significantly improved remote sensing (RS) multimodal understanding. Language-conditioned segmentation is crucial for fine-grained target understanding in Unmanned Aerial Vehicle (UAV) videos. However, this task remains challenging due to the prevalence of small, visually ambiguous targets and dynamic aerial perspectives. In this paper, we propose SkyVLaM, a multimodal large language model for UAV video understanding. SkyVLaM constructs sparse tokens directly from patch-level video representations through a temporal basis perceiver, regularizes the sparse basis to encourage complementary temporal cues, and adaptively selects a temporally coherent dense segment for high-resolution inspection. The resulting sparse and dense tokens are jointly processed by a large language model for query-conditioned segmentation. We further build SkyVid, consisting of SkyVid-VGCG and SkyVid-RVOS for video grounded conversation generation and referring video object segmentation, respectively. SkyVid contains 101 videos, 33.6K frames, and 1.53M pixel-level object instances. Experiments show that SkyVLaM provides a more effective allocation of the visual token budget and improves language-conditioned video segmentation in UAV scenarios.
CommentsAccepted by WAICA 2026