arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LongEarth-R1:用于长时序地球观测推理的视觉语言模型基准测试与对齐

LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning

Yupan Ding, Jing Xiao, Zhenyuan Zhang, Chaofeng Chen, Liang Liao, Gui-Song Xia, Mi Wang

arXiv 2608.13344首次发表:更新:

AI 中文总结

该研究推出长时序地球观测推理基准LongEarth-Bench,开发LongEarth-R1模型,在全部12项长序列遥感推理任务中表现最优,同时在标准遥感基准上保持竞争力。

AI 中文摘要

长时序地球观测推理要求模型能够梳理多阶段地理演化、定位空间变化、检测时间异常,并从扩展的图像序列中进行未来推断。然而,现有遥感视觉语言模型主要聚焦于孤立图像、图像对或短序列,限制了其对相关帧和区域的可靠定位。我们推出LongEarth-Bench,该基准包含约12万个问答样本,源自11.7万张独特图像。其序列平均长度为15.14帧,最长可达30帧,涵盖演化总结、空间推理、异常识别及逻辑预测等12项任务。另有3万个样本的子集提供结构化推理轨迹,将关键帧与变化区域关联至最终答案。我们通过结合显式序列标识符与结构化思维链监督的监督微调开发了LongEarth。基于LongEarth,LongEarth-R1应用了包含格式、时间及空间奖励的组相对策略优化。LongEarth-R1在全部12项长序列任务中取得最佳结果,同时在标准遥感基准上保持竞争力。

英文摘要

Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image sequences. However, existing remote sensing vision-language models mainly focus on isolated images, image pairs, or short sequences, limiting reliable grounding in the relevant frames and regions. We introduce LongEarth-Bench, a benchmark containing approximately 120k question-answering samples derived from 117k unique images. Its sequences average 15.14 frames and extend to 30 frames, covering 12 tasks across evolution summarization, spatial reasoning, anomaly identification, and logical prediction. A 30k-sample subset further provides structured reasoning traces linking key frames and changed regions to final answers. We develop LongEarth through supervised fine-tuning with explicit sequence identifiers and structured chain-of-thought supervision. Building on LongEarth, LongEarth-R1 applies group relative policy optimization with format, temporal, and spatial rewards. LongEarth-R1 achieves the best results on all 12 long-sequence tasks while remaining competitive on standard remote sensing benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑