arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种基于仿真锚定的智能体视觉语言模型框架,用于野火监测与报告

A Simulation-Grounded Agentic VLM Framework for Wildfire Monitoring and Reporting

Duowen Chen, Yuchen Sun, Zhiqi Li, Yuxuan Liao, Sinan Wang, Bart van Bloemen Waanders, Bo Zhu

arXiv 2610.02451首次发表:更新:

发表机构

Georgia Institute of Technology; Sandia National Laboratories(佐治亚理工学院; 桑迪亚国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出基于仿真锚定的VLM框架,自动将2D野火仿真转为带标签视频,结合多智能体记忆检索与推理,实现高效野火监测报告,显著提升标签和报告字段准确率。

AI 中文摘要

有效的野火监测需要将视觉证据与物理火灾动态相关联,然而具有同步物理标注的真实视频十分稀缺,且高保真三维仿真成本高昂。我们提出了一种基于仿真锚定的视觉语言模型(VLM)框架,该框架可自动将二维野火仿真转换为带标签的视频片段。固定的Blender映射可生成与仿真器地形、燃料布局、火灾活动和风场线索对齐的低细节三维代理模型;可控视频生成则提供更丰富的外观。这些代理模型是中间表示,而非精细渲染的最终场景。生成的视频和仿真器标签构成可复用的多模态记忆,供一个免训练的多智能体VLM系统使用,该系统可检索参考片段、协调视觉与基于记忆的预测,并生成结构化的野火报告。在保留的生成片段上,视频记忆实现了51.5%的精确四标签准确率,而直接VLM查询为22.6%,纯文本记忆为16-17%;完整系统在六个源自仿真器的报告字段上达到了77.3%的准确率。组件消融实验、跨生成器测试以及三次真实无人机评估,分别检验了检索、报告、生成器变化和可观测监测任务。该框架在真实物理标注稀缺的情况下,将自动仿真到代理转换与基于记忆的VLM推理相连接。

英文摘要

Effective wildfire monitoring requires relating visual evidence to physical fire dynamics, yet real videos with synchronized physical annotations are scarce and high-fidelity 3D simulation is costly. We present a simulation-grounded vision-language model (VLM) framework that automatically converts 2D wildfire simulations into labeled video episodes. A fixed Blender mapping produces low-detail 3D proxies aligned with simulator terrain, fuel layout, fire activity, and wind cues; controllable video generation supplies richer appearance. The proxies are intermediate representations rather than finely rendered final scenes. Generated videos and simulator labels form reusable multimodal memory for a training-free multi-agent VLM system that retrieves reference episodes, reconciles visual and memory-based predictions, and produces structured wildfire reports. On held-out generated episodes, video memory achieves 51.5% exact four-tag accuracy, compared with 22.6% for direct VLM querying and 16-17% for text-only memory; the complete system achieves 77.3% accuracy on six simulator-derived report fields. Component ablations, cross-generator tests, and three real-UAV evaluations assess retrieval, reporting, generator changes, and observable monitoring tasks. The framework connects automatic simulation-to-proxy conversion with memory-based VLM reasoning under scarce real-world physical annotations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑