arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于流式航空视频中小目标理解的内存增强多模态大语言模型

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos

Penglei Sun, Yehua Huang, Zhuoli Tao, Xiang Li, Runwei Guan, Yaoxian Song, Kaiyong Zhao, Henghui Ding, Bo Han, Yang Yang, Xiaowen Chu

arXiv 2607.19857首次发表:更新:

发表机构

The Hong Kong University of Science and Technology (Guangzhou); University of Freiburg; Hangzhou City University; XGRIDS; Fudan University; Hong Kong Baptist University(香港科技大学(广州); 弗莱堡大学; 杭州城市大学; XGRIDS公司; 复旦大学; 香港浸会大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对流式航空视频中小目标理解难题,提出像素级开放词汇数据集DroneEyes,以及含语义感知令牌路由器和分层内存库的多模态大语言模型SkyAnchor,从数据和方法角度应对挑战。

AI 中文摘要

语言引导的航空感知旨在理解复杂无人机场景中用户指定的微小目标。在实际无人机部署中,无人机飞行时必须做出响应,这种感知以在线流式方式运行,帧按顺序到达,模型对每个帧做出响应而无法访问未来帧。将当前多模态大语言模型应用于此设置存在两个挑战。一是空中看到的目标通常很小,现有模型的视觉压缩对所有区域一视同仁,丢弃了细粒度细节;二是理解连续流需要过去帧的上下文,但在资源受限的机载硬件上保留整个历史是不可行的,而丢弃则会导致目标漂移或消失。我们从数据和方法角度应对这些挑战。从数据角度,我们提出了DroneEyes,这是第一个用于微小航空目标的像素级和开放词汇引用分割数据集。从方法角度,我们提出了SkyAnchor,一种针对上述挑战有两种设计的多模态大语言模型:一种语义感知令牌路由器,在减少视觉令牌预算的情况下保留小目标;一种分层内存库,在流上保持对目标的持续理解。

英文摘要

Language-guided aerial perception aims to understand user-specified tiny targets in complex unmanned aerial vehicle (UAV) scenes. In real UAV deployment, the UAV must respond while it flies, so such perception runs in an online streaming manner, where frames arrive sequentially and the model responds to each one without access to future frames. However, applying current Multimodal Large Language Models (MLLMs) to this setting raises two challenges. First, targets viewed from the air are often tiny, yet the visual compression in existing MLLMs treats all regions equally and discards their fine-grained details. Second, understanding a continuous stream requires past-frame context, yet retaining the entire history is infeasible on resource-constrained onboard hardware, whereas discarding it causes the target to drift or disappear. We address the tiny object and streaming challenges from both data and method perspectives. From the data perspective, we present \textbf{DroneEyes}, the \textbf{first} pixel-level and open-vocabulary referring-segmentation dataset for tiny aerial targets, comprising $2,140$ high-definition videos and $176,623$ pairs across Object Description and Referring Expression tasks, with dense per-frame masks. From the method perspective, we propose \textbf{SkyAnchor}, an MLLM with two designs to the above challenges: a Semantics-Aware Token Router that preserves small-target under a reduced visual-token budget, and a Hierarchical Memory Bank that keeps the target consistently understood on streams.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑