arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GeoForge:用于对地观测推理的非参数自进化智能体

GeoForge: Non-Parametric Self-Evolving Agents for Earth-Observation Reasoning

Xin Xiao, Jiang Zhong, Junnan Zhu, Yingchao Feng, Peijin Wang, Yidan Zhang, Kaiwen Wei

arXiv 2608.10494首次发表:更新:

AI 中文总结

GeoForge是一种无需训练的自进化框架,通过结构化非参数执行状态及多记忆机制提升EO智能体的规划与推理能力,可改善任务准确率、工具轨迹质量并减少推理错误。

AI 中文摘要

对地观测(EO)智能体构建科学有效的工具工作流,并以当前地理空间证据为依据得出结论。这一任务颇具挑战性,因为EO工作流受限于感知语义、产品依赖关系、时空兼容性及参数要求。现有智能体通常会针对每个查询在广阔的操作空间中搜索,而近期的自进化系统未能将异构EO轨迹充分组织为可跨不同决策层级复用的知识。为解决该问题,我们提出GeoForge,这是一种无需训练的自进化框架,可将已完成的轨迹转换为结构化非参数执行状态。GeoForge会根据感知上下文约束操作空间,再从三个互补记忆中检索任务条件先验:工作流图记忆捕获全局操作顺序,动作级经验提供局部修正,适配技能标准操作流程保留程序及数据约束。检索到的先验指导工具执行,而当前观测结果仍是最终答案的依据。每项任务完成后,安全门控蒸馏过程会将已落地轨迹转换为可复用的执行知识,供后续检索使用。这种执行、蒸馏与复用的循环无需更新骨干大语言模型(LLM)即可改进规划。在多个地理空间基准上开展的实验表明,GeoForge在各类LLM骨干上均能持续提升任务准确率与工具使用轨迹质量,同时大幅降低多数LLM的工具规划与推理错误。

英文摘要

Earth observation (EO) agents construct scientifically valid tool workflows and ground their conclusions in current geospatial evidence. This is challenging because EO workflows are constrained by sensing semantics, product dependencies, spatial and temporal compatibility, and parameter requirements. Existing agents often search a broad operation space for each query, while recent self-evolving systems do not fully organize heterogeneous EO trajectories into reusable knowledge across different decision levels. To solve this problem, we present GeoForge, a training-free, self-evolving framework that transforms completed trajectories into a structured nonparametric execution state. GeoForge constrains the operation space according to the sensing context, then retrieves a task-conditioned prior from three complementary memories. Workflow Graph Memory captures global operation order, Action-Level Experiences provide local corrections, and the Adapted Skill Standard Operating Procedure preserves procedural and data constraints. The retrieved prior guides tool execution, while current observations remain the basis of the final answer. After each task, a safety-gated distillation process converts grounded trajectories into reusable execution knowledge for future retrieval. This execution, distillation, and reuse loop improves planning without updating the backbone LLM. Experiments on multiple geospatial benchmarks demonstrate that GeoForge consistently improves both task accuracy and tool-use trajectory quality across diverse LLM backbones, while substantially reducing tool-planning and reasoning errors for most LLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑