发表机构
Tianjin University; Aerospace Information Research Institute, Chinese Academy of Sciences(天津大学; 中国科学院空天信息创新研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RSure-Agent通过可验证观测协议和任务-工具可靠性先验,验证遥感智能体的工具观测并限制错误传播,在多个基准上显著降低错误传播率并提升准确率。
AI 中文摘要
遥感智能体依赖感知、测量和栅格分析工具来解决地球观测任务。我们将它们关于地物的判断和定量结果称为工具观测。然而,这些观测存在很大的不确定性,即使工具执行成功,也可能是错误的。当智能体接受错误的观测时,错误可能会在后续推理中传播并导致任务失败。我们分析了三个遥感智能体基准上的1,229条执行轨迹。在每个基准上,至少88.1%的任务依赖于工具观测。在这些任务中,至少22.7%的任务在工具成功执行的情况下仍包含错误的观测。在每个基准上,这些错误在至少82.0%受影响的任务中传播到最终答案。为了解决这个问题,我们提出了RSure-Agent,一个用于验证工具观测并限制错误传播的框架。我们引入了一个可验证的观测协议,要求工具返回过程证据供智能体验证其观测。我们还从离线任务反馈中构建了一个任务-工具可靠性先验。该先验总结了每个工具配置在任务类型上的历史表现,并为验证提供了任务特定的参考。利用过程证据和这一先验,RSure-Agent决定是接受观测、请求额外证据还是拒绝它。我们在EarthBench、ThinkGeo、TerraLogic和CHOICE-420上评估了RSure-Agent。在三个智能体基准上,相对于禁用验证和先验的基础配置,RSure-Agent将错误传播率降低了21.3到25.9个百分点。在CHOICE-420上,它在11个骨干模型上平均将直接回答的整体准确率提高了5.71个百分点。在EarthBench上,相对于Earth-Agent,它将工具调用比率降低了25.9%。
英文摘要
Remote sensing agents rely on perception, measurement, and raster analysis tools to solve Earth observation tasks. We refer to their judgments and quantitative results about ground objects as tool observations. However, these observations are subject to substantial uncertainty and may be incorrect even when the tools execute successfully. When agents accept incorrect observations, the errors can propagate through subsequent reasoning and cause task failure. We analyze 1,229 execution trajectories across three remote sensing agent benchmarks. On each benchmark, at least 88.1% of tasks depend on tool observations. Among these tasks, at least 22.7% contain incorrect observations despite successful tool execution. These errors propagate to the final answer in at least 82.0% of affected tasks on each benchmark. To address this problem, we propose RSure-Agent, a framework for verifying tool observations and limiting error propagation. We introduce a verifiable observation protocol that requires tools to return process evidence for the agent to verify their observations. We also construct a task-tool reliability prior from offline task feedback. The prior summarizes each tool configuration's past performance across task types and provides a task-specific reference for verification. Using process evidence and this prior, RSure-Agent decides whether to accept an observation, request additional evidence, or reject it. We evaluate RSure-Agent on EarthBench, ThinkGeo, TerraLogic, and CHOICE-420. Across the three agent benchmarks, RSure-Agent reduces the error propagation rate by 21.3 to 25.9 percentage points relative to the base configuration with verification and the prior disabled. On CHOICE-420, it improves overall accuracy over direct answering by 5.71 percentage points on average across 11 backbone models. On EarthBench, it reduces the tool-call ratio by 25.9% relative to Earth-Agent.
CommentsThe demo is available at https://github.com/airs101/RSure-Agent (code will be released for further research)