arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15768cs.CVcs.AI

GeoChrono:遥感中长期时间理解的基准测试与重新思考

GeoChrono: Benchmarking and Rethinking Long-Term Temporal Understanding in Remote Sensing

Yujie Li, Jiancheng Pan, Zhiwei Wei, Jiuniu Wang, Mugen Peng, Wenjia Xu

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对遥感中长期时间理解缺乏系统评估的问题,引入ChronoBench基准。基于评估结果提出GeoChrono模型,设计相关编码器和压缩器并构建训练数据集,该模型在基准测试中性能领先,压缩器有效减少视觉令牌。

中文摘要 AI 辅助

遥感为观察地球长期表面演化提供了独特视角,要求模型不仅能感知孤立时刻的土地覆盖,还能追踪变化、记忆演化历史并进行时空推理。现有研究缺乏系统评估。为此引入ChronoBench基准,将任务分解为四个认知水平,包含12个子任务和大量问答对。评估发现主流语言模型落后于人类专家,长期记忆是关键瓶颈。在此基础上提出GeoChrono模型,设计了时间轨迹编码器和粗到细令牌压缩器,并构建训练数据集。GeoChrono在基准测试中取得领先,压缩器减少视觉令牌同时保持高准确率。代码和数据将在指定网址提供。

英文摘要

Remote sensing offers an unparalleled vantage point for observing the Earth's long-term surface evolution, yet it demands that a model not only perceive land cover at isolated moments, but also track changes, memorize evolution histories, and reason across time and space. However, existing studies lack a systematic evaluation that dissects these distinct competencies. To fill this gap, we introduce ChronoBench, a multidimensional benchmark that decomposes this task into four progressive cognitive levels (i.e., Land Cover Perception, Temporal Recognition, Long-Term Memory, and Spatio-Temporal Reasoning). The ChronoBench comprises 12 sub-tasks and 17,689 rigorously validated QA (Question-Answer) pairs. Extensive evaluations reveal that mainstream MLLMs fall drastically behind human experts, with Long-Term Memory emerging as the most critical bottleneck. Motivated by this finding, we further propose GeoChrono, an MLLM with enhanced capabilities for tracing, memorizing, and reasoning about long-term geographic evolution. Leveraging the physical prior that geographic parcels remain spatially fixed while their semantics evolve, we design a Temporal Trajectory Encoder~(TempEnc) that constructs per-location temporal trajectories for dedicated land cover evolution modeling, and we introduce a Coarse-to-Fine Token Compressor~(C2FComp) that adaptively preserves dynamic regions while compressing the static background. To support training, we also construct ChronoInstruct, a 104K-sample instruction-tuning dataset spanning all competency levels for training. GeoChrono achieves state-of-the-art performance on ChronoBench, surpassing the leading commercial MLLMs by over 20%, while C2FComp reduces visual tokens by over 56% while retaining GeoChrono's 94.6% performance. The code and data will be available at https://github.com/IntelliSensing/GeoChrono

发表机构

  • Beijing University of Posts and Telecommunications(北京邮电大学)
  • Tsinghua University(清华大学)
  • Hunan Normal University(湖南师范大学)
  • City University of Hong Kong(香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑