arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RS-Claw-Evolution:面向长时程任务中轻量级遥感智能体的环境反馈驱动进化

RS-Claw-Evolution: Environment-Feedback-Driven Evolution for Lightweight Remote Sensing Agents in Long-Horizon Tasks

Kai Ouyang, Dongyang Hou, Liangtian Liu, Zeyuan Wang, Ziyu Li, Chengfu Liu, Zichao Tang, Xuezhi Cui, Shengwu Ouyang, Wentao Yang, Hanwen Yu, Haifeng Li

arXiv 2609.22258首次发表:更新:

发表机构

Central South University; Hunan University of Science and Technology; University of Electronic Science and Technology of China(中南大学; 湖南科技大学; 电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出环境反馈驱动的RS-Claw-Evolution框架,通过交互、经验与决策三阶段进化,使轻量级Qwen3-4B智能体在长时程遥感任务中超越更大模型,接近GPT-5。

AI 中文摘要

大型语言模型驱动的遥感(RS)智能体为自动化地理空间分析提供了一种有前景的方法。然而,基于紧凑语言模型的轻量级遥感智能体在多步交互式任务中表现不佳,原因在于长时程状态的丢失、环境反馈利用效率低下以及优化信号稀疏。我们提出了RS-Claw-Evolution,一个环境反馈驱动的框架,通过三个阶段逐步改进轻量级智能体。交互进化使用可执行代码来控制观察、维护中间状态并减少上下文冗余。经验进化将失败感知的轨迹生成与错误回合掩码相结合,从信息丰富的失败恢复经验中学习,而不模仿错误动作。决策进化使用具有多维环境奖励和回合级优势保护(turn-level advantage protection)的强化学习来优化工具使用行为,并改善长序列中的信用分配。在Earth-Bench上,优化后的基于Qwen3-4B的智能体在自主规划模式下达到了65.9%的准确率,优于未训练的Qwen3-32B基线(43.8%)和DeepSeek-V3.1(60.8%),并接近GPT-5(71.6%)。这些结果表明,从环境反馈中学习可以改进轻量级智能体,并缩小它们与更大模型在长时程遥感任务中的性能差距。

英文摘要

Large language model-driven remote sensing (RS) agents offer a promising approach to automating geospatial analysis. However, lightweight RS agents based on compact language models struggle with multi-step interactive tasks due to loss of long-horizon states, inefficient environmental feedback utilization, and sparse optimization signals. We propose RS-Claw-Evolution, an environment-feedback-driven framework that progressively improves lightweight agents through three stages. Interaction evolution uses executable code to control observations, maintain intermediate states, and reduce context redundancy. Experience evolution combines failure-aware trajectory generation with error-turn masking to learn from informative failure-recovery experiences without imitating faulty actions. Decision evolution uses reinforcement learning with multi-dimensional environment rewards and turn-level advantage protection to optimize tool-use behaviors and improve credit assignment in long sequences. On Earth-Bench, the optimized Qwen3-4B-based agent achieves 65.9% accuracy in Autonomous Planning mode, outperforming the untrained Qwen3-32B baseline (43.8%) and DeepSeek-V3.1 (60.8%), while approaching GPT-5 (71.6%). These results demonstrate that learning from environmental feedback can improve lightweight agents and narrow their performance gap with larger models in long-horizon RS tasks.

Comments30 pages, 5 figures, including supplementary material

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑