先总结,后下载:面向带宽高效的地球观测的机载视觉语言模型
Summarize First, Download Later: Onboard VLMs for Bandwidth-Efficient Earth Observation
浏览论文内容
中文总结 AI 辅助
针对地球观测卫星下行链路带宽瓶颈,提出“先总结,后下载”范式,通过机载VLM生成摘要、地面VQA验证后按需下载全分辨率图像,可降带宽消耗并加速时间敏感任务的洞察时间。
中文摘要 AI 辅助
现代地球观测(EO)卫星搭载的传感器愈发先进,可生成海量高分辨率多光谱数据,但下行链路容量仍是关键瓶颈,常导致显著延迟或在有限通信窗口内丢失宝贵观测数据。我们提出一种“先总结,后下载”范式,利用机载边缘计算与视觉语言模型(VLM)的最新进展。系统并非盲目下行原始图像,而是遵循三阶段交互协议:卫星先传输由量化机载VLM生成的简明自然语言摘要;地面操作人员随后提出针对性视觉问答(VQA)查询,验证场景相关性(如野火或海上异常);仅在确认关键信息时才下载全分辨率图像。这将下行链路从被动批量传输转变为主动、语义感知的对话。我们在资源受限的NVIDIA Jetson平台上实现并评估该系统,对各类遥感场景的实验表明,所提策略大幅降低带宽消耗,同时加快时间敏感任务的洞察时间。
英文摘要
Modern Earth observation (EO) satellites carry increasingly advanced sensors that produce vast volumes of high-resolution, multispectral data, yet downlink capacity remains a critical bottleneck -- often causing significant latency or the loss of valuable observations within limited contact windows. We propose a "Summarize First, Download Later" paradigm that exploits recent advances in onboard edge computing and Vision-Language Models (VLMs). Rather than indiscriminately downlinking raw imagery, the system follows a three-phase interaction protocol: the satellite first transmits concise natural language summaries generated by a quantized onboard VLM; ground operators then issue targeted Visual Question Answering (VQA) queries to verify scene relevance (e.g., wildfires or maritime anomalies); and full-resolution images are downloaded only when critical information is confirmed. This transforms the downlink from passive bulk transfer into an active, semantics-aware dialogue. We implement and evaluate the system on a resource-constrained NVIDIA Jetson platform, and experiments on diverse remote sensing scenes show that the proposed strategy substantially reduces bandwidth consumption while accelerating time-to-insight for time-sensitive missions.