arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31234cs.CV

WeaveAgent:面向超高分辨率遥感影像的两阶段工具路由智能体

WeaveAgent: A Two-Stage Tool-Routing Agent for Ultra-High-Resolution Remote Sensing Imagery

  • National University of Defense Technology(国防科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Zhongyu Pang

AI总结:

WeaveAgent提出两阶段工具路由方法,将路由与视觉感知解耦,通过SFT和GRPO训练,显著提升外在路由准确率,并揭示了观察加载和两轮SFT对答案准确率的贡献。

AI中文摘要:

问题。具有模糊用户意图的超高分辨率(UHR)遥感存在两个瓶颈:视觉标记成本高昂,且工具调用必须格式可靠(预训练模型零样本时不会发出任何工具调用)。方法。WeaveAgent,一个两阶段工具路由智能体,将路由与视觉感知解耦。阶段A以路由优先:发射是经过训练的,而非诱发的。阶段B有条件地执行:内在查询进入视觉回答(全场景缩略图;一个WeaveEarth风格的证据板作为可选固定预算、约5k标记的压缩接口);外在查询在原始全分辨率影像上执行工具调用,并在第二轮观察掩蔽轮次中根据工具观察结果进行回答。训练:对齐SFT,然后在奖励R_WA2下进行GRPO。结果。对齐SFT将外在路由从0%提升至80.75%(323/400);GRPO抑制了9次内在误发射,而工具选择保持不变。训练后的2B系统整体上未超越零样本8B基线(0.263对0.250),这是一个诊断性贡献。Oracle归因分离了两种修复成分:将观察加载到上下文中,在无标记跨模态返回下将外在答案准确率从0.025提升至0.425,而两轮SFT阶段进一步增加+9.3个百分点至0.518,代价是较小的路由成本。+/-图像消融显示发射抑制具有视觉基础,查询注册矩阵显示LLM重写的查询使训练后的检查点损失2-11个百分点。范围。所有训练和评估均使用5,000/3,273/1,000条记录的VagueUHR语料库(600条内在+400条需要工具;基础种子用于合成,不用于优化)。单遍证据构建在RTX 4090上每张图像运行7.31秒。代码、数据和评估协议将发布。

英文摘要:

Problem. Ultra-high-resolution (UHR) remote sensing with vague user intents has two bottlenecks: visual tokens are expensive, and tool calling must be format-reliable (pretrained models emit zero tool calls zero-shot). Method. WeaveAgent, a two-stage tool-routing agent, decouples routing from visual perception. Stage A is routing-first: emission is trained, not elicited. Stage B executes conditionally: intrinsic queries enter visual answering (full-scene thumbnail; a WeaveEarth-style evidence board as an optional fixed-budget, approx. 5k-token compression interface); extrinsic queries execute tool call on original full-resolution imagery, answering from tool observations in a second, observation-masked round. Training: alignment SFT, then GRPO under reward R_WA2. Results. Alignment SFT lifts extrinsic routing from 0% to 80.75% (323/400); GRPO suppresses 9 intrinsic mis-emissions while tool selection is unchanged. The trained 2B system does not beat the zero-shot 8B baseline overall (0.263 vs. 0.250), a diagnostic contribution. Oracle attribution separates two repair ingredients: loading the observation into context lifts extrinsic answer accuracy from 0.025 to 0.425 under marker-free cross-mode returns, and the two-turn SFT stage adds a further +9.3 points to 0.518 at a small routing cost. A +/- image ablation shows emission suppression is visually grounded, and a query-register matrix shows LLM-rewritten queries cost trained checkpoints 2-11 points. Scope. All training and evaluation use the 5,000 / 3,273 / 1,000-record VagueUHR corpus (600 intrinsic + 400 tool-requiring; the base seeds synthesis and is not used for optimization). Single-pass evidence construction runs at 7.31 s per image on an RTX 4090. Code, data, and evaluation protocols will be released.

↑