发表机构
Shanghai Jiao Tong University; Shanghai Artificial Intelligence Laboratory; Tianjin University; Harbin Institute of Technology; Tsinghua Shenzhen International Graduate School; Tsinghua University; Sun Yat-Sen University(上海交通大学; 上海人工智能实验室; 天津大学; 哈尔滨工业大学; 清华大学深圳国际研究生院; 清华大学; 中山大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Earth-Agent-Pro提出执行自适应的计划与执行框架,利用专家技能和结构化记忆实现全链路地球观测,在Earth-Bench-Pro上显著超越ReAct,并提升小模型性能。
AI 中文摘要
真实世界的地球观测(EO)智能体必须将高层科学问题转化为可执行的工作流,以获取观测数据、准备数据、执行领域计算,并从运行时证据中得出结论。现有地球观测智能体通常从已提供的观测数据开始,而基准测试通常提供准备好的输入或候选答案,使得全链路的开放世界地球观测执行在很大程度上未经测试。我们提出了Earth-Agent-Pro,一个执行自适应的计划与执行(Plan-and-Execute)框架,使用专家编写的技能来约束规划和运行时工具使用。以工作流为中心的结构化记忆记录了计划步骤、已接受的证据及其依赖关系,使得当运行时证据使某一步骤失效时,仅需修复受影响的工作流后缀。独立的大语言模型适配器使用序列级监督微调进行规划器工作流组合,以及使用节点级组相对策略优化和局部可验证奖励进行执行器工具参数接地。Earth-Bench-Pro将248个专家策划的任务核心实例化为三个匹配机制下的744个问题。其248个开放世界执行问题涵盖RGB图像、光谱观测和遥感产品,将高层请求与运行时数据需求、可执行轨迹以及基于执行证据的开放式答案配对。在共享GPT-5骨干下,Earth-Agent-Pro实现了66.13%的LLM-as-Judge准确率,在该指标上超过ReAct 20.95个百分点,在工具顺序(Tools-In-Order)上超过24.44个百分点。联合适配器调优将Qwen3.5-9B的LLM-as-Judge准确率从38.31%提升至50.00%,比未调优配置提高了11.69个百分点。仅规划评估和使用参考工作流的执行表明,适配器分别改善了工作流组合和参数接地。代码和数据集将很快发布。
英文摘要
Real-world Earth observation (EO) agents must translate high-level scientific questions into executable workflows to acquire observations, prepare data, perform domain computations, and derive conclusions from runtime evidence. Existing EO agents typically start from supplied observations, while benchmarks typically provide prepared inputs or candidate answers, leaving full-chain open-world EO execution largely untested. We present Earth-Agent-Pro, an execution-adaptive Plan-and-Execute framework using expert-authored skills to constrain planning and runtime tool use. Workflow-centered structured memory records planned steps, accepted evidence, and their dependencies, enabling repair of only the affected workflow suffix when runtime evidence invalidates a step. Separate large language model adapters use sequence-level supervised fine-tuning for planner workflow composition and node-level group relative policy optimization with locally verifiable rewards for executor tool-argument grounding. Earth-Bench-Pro instantiates 248 expert-curated task cores as 744 questions under three matched regimes. Its 248 Open-World Execution questions span RGB imagery, spectral observations, and remote sensing products, pairing high-level requests with runtime data requirements, executable trajectories, and open-ended answers grounded in execution evidence. With a shared GPT-5 backbone, Earth-Agent-Pro achieves 66.13% LLM-as-Judge accuracy, exceeding ReAct by 20.95 points in this metric and 24.44 points in Tools-In-Order. Joint adapter tuning raises Qwen3.5-9B LLM-as-Judge accuracy from 38.31% to 50.00%, an 11.69-point gain over the untuned configuration. Planning-only evaluation and execution with the reference workflow show that the adapters improve workflow composition and argument grounding, respectively. Code and datasets will be released soon.
Comments18 pages, 8 figures. The code and datasets of this work will be released soon