arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24772cs.AIcs.CL

RSMeM:用于遥感智能体的知识增强记忆进化及系统评估

RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic Evaluation

Bingxian Wu, Yu Zhang, Zonghao Guo, Tang Liu, Chen Qian, Yuxiang Lu, Xingbo Du, Yanghao Li, Yidan Zhang, Chi Chen, Ling Yao, Maosong Sun

AI总结:

研究针对现有遥感智能体问题,提出RSMeM机制,通过分层知识基础和失败感知经验提炼两个组件,迭代吸收领域知识转化为执行经验,经实验验证能提升工具使用性能和答案质量,具有强大知识密度。

AI中文摘要:

地球科学研究需要复杂分析和领域专业知识,遥感观测是关键基础。现有基于通用语言模型构建的遥感智能体在很大程度上与领域无关,工作流程脆弱且易出错,失败经验很少整合为可复用经验。为此引入RSMeM,一种知识增强记忆进化机制,用预提炼领域知识引导遥感智能体并迭代整合在线经验以实现稳健多步工具执行。它由分层知识基础和失败感知经验提炼两个组件构成。通过迭代使用这两个过程,智能体可吸收任务级领域知识并转化为实例级执行经验。在EarthBench上的大量实验表明,RSMeM持续提升工具使用性能和端到端答案质量,并展示出强大的知识密度。

英文摘要:

Geoscience research requires complex analysis and domain expertise, with remote sensing (RS) observations as a key foundation. However, existing RS agents built on general-purpose LLMs remain largely domain-agnostic, resulting in brittle and error-prone workflows. Moreover, these failures are seldom consolidated into a reusable experience for subsequent analyses. To address this issue, we introduce RSMeM, a knowledge-enhanced memory evolution mechanism that bootstraps RS agents with pre-distilled domain knowledge and iteratively integrates online experience for robust multi-step tool execution. RSMeM is composed of two components: (i) Hierarchical Knowledge Grounding, which performs taxonomy-aware retrieval over a hierarchical domain corpus to guide planning and tool selection; and (ii) Failure-Aware Experience Refinement, which distills failure-annotated tool-use traces into reusable constraints for next-round tool execution. By iteratively employing these two processes, RS agents can evolve to absorb task-level domain knowledge and effectively translate it into instance-level execution experience. Extensive experiments on EarthBench demonstrate that RSMeM consistently improves tool-use performance and end-to-end answer across a diverse set of LLM backbones. Notably, RSMeM achieves a 6% accuracy improvement on DeepSeek-V3.2 with less than 1% additional experience tokens, demonstrating the strong knowledge density of our distilled experience.

补充信息

↑