遥感中智能体变化视觉问答的选择性工具使用
Selective Tool Use for Agentic Change Visual Question Answering in Remote Sensing
浏览论文内容
中文总结 AI 辅助
针对遥感变化视觉问答中VLM不可靠问题,提出选择性工具使用框架,通过调用确定性工具获取语义证据,显著提升准确率,并验证了语义图质量的影响。
中文摘要 AI 辅助
变化视觉问答(Change VQA)要求理解双时相遥感图像中的语义变化。尽管视觉语言模型(VLMs)在此任务上展现出有前景的性能,但在回答需要明确转换统计、面积测量或空间信息的问题时,它们仍不可靠。为解决此局限,我们提出一个选择性工具使用框架,其中单个VLM要么直接回答,要么调用确定性变化分析工具以获取针对问题的证据。具体而言,所选工具操作于双时相语义图,并返回结构化观察结果,同一VLM利用该结果生成最终答案。为支持此框架,我们构建了CDVQA的工具增强扩展,涵盖八个问题族和三个用于转换、空间及时间分析的工具。工具使用监督和观察结果自动从原始语义标注中派生,无需额外人工标注。随后,我们使用低秩适配(LoRA)适配Qwen3.5-4B,以联合学习直接回答、工具调用和基于证据的回答。在7,164个测试问题上的实验表明,使用参考语义图的选择性工具使用将整体准确率从73.77%提升至88.79%,平均族准确率从69.11%提升至89.65%。当语义图自动预测时,该框架达到77.47%的整体准确率和75.06%的平均族准确率。这些结果证明了针对问题的语义证据对Change VQA的益处,同时强调了语义预测质量对最终性能的影响。代码和工具增强标注将在此https URL公开提供。
英文摘要
Change visual question answering (Change VQA) requires understanding semantic changes across bi-temporal remote sensing images. Although vision language models (VLMs) have shown promising performance on this task, they remain unreliable when answering questions that require explicit transition statistics, area measurements, or spatial information. To address this limitation, we propose a selective tool use framework in which a single VLM either answers directly or invokes a deterministic change analysis tool to obtain question specific evidence. Specifically, the selected tool operates on bi-temporal semantic maps and returns a structured observation, which the same VLM uses to generate its final answer. To support this framework, we construct a tool augmented extension of CDVQA covering eight question families and three tools for transition, spatial, and temporal analysis. Tool use supervision and observations are derived automatically from the original semantic annotations, without additional manual labeling. We then adapt Qwen3.5-4B using Low Rank Adaptation (LoRA) to jointly learn direct answering, tool invocation, and evidence conditioned answering. Experiments on 7,164 test questions show that selective tool use with reference semantic maps improves overall accuracy from 73.77% to 88.79% and average family accuracy from 69.11% to 89.65%. When the semantic maps are predicted automatically, the framework achieves 77.47% overall accuracy and 75.06% average family accuracy. These results demonstrate the benefit of question-specific semantic evidence for Change VQA, while highlighting the influence of semantic prediction quality on the resulting performance. Code and tool-augmented annotations will be made publicly available at https://github.com/yakoubbazi/ToolChangeVQA.
发表机构
- King Saud University(沙特国王大学)
机构由 AI 辅助整理,请以论文原文为准。