arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于检索增强推理、物体感知与损伤分析的集成多模态AI系统

Integrated Multimodal AI System for Retrieval-Augmented Reasoning, Object Sensing, and Damage Analysis

Kalelo Dukuray, Israel Pina, Evan Perez, Erika Ardiles-Cruz, Jie Wei

arXiv 2608.08935首次发表:更新:

发表机构

City College of New York; Air Force Research Lab(纽约城市学院; 空军研究实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出整合RAG、热谱感知等技术的多模态AI系统,用于损伤评估,通过多模块对比实验验证了其在提升推理准确性、鲁棒性及跨场景检测方面的优势。

AI 中文摘要

本研究提出一种用于损伤评估的统一多模态AI系统,整合了检索增强生成(RAG)模型、热谱感知、视觉基础模型流水线及探索性无线信号感知。研发了一个RAG组件,将本地部署的语言模型锚定在项目特定文档中,包括专门的损伤等级分类标准,以缓解推理过程中的幻觉问题。与静态少样本提示的对照实验表明,动态检索可提升锚定效果与事实一致性。我们进一步对比了基于向量的RAG与通过实体-关系抽取构建的知识图谱变体,结果显示,基于图谱的检索在需要跨文档推理的损伤评估查询中能产生更强的响应,这推动了混合密集、稀疏与图谱感知检索的发展。为解决EO(光电)图像在恶劣光照与天气条件下的局限性,采用红外(IR)/热感知进行物体检测与分割,我们的检测器生成候选检测结果,实现了对广泛物体的更优分割。成对的红外与可见光谱跟踪实验揭示了失效模式,这推动了用于鲁棒物体检测与损伤分析的多模态融合技术的发展。利用视觉基础模型与视觉-语言模型生成合成损伤图像并高精度分类损伤严重程度,支持下游损伤评估模型的训练与验证。最后,探索性无线感知展示了在EO与IR感知失效的场景下,检测存在性、运动及事件后环境变化的潜力。

英文摘要

This work presents a unified multimodal AI system for damage assessment that integrates retrieval-augmented generation (RAG) models, thermal spectrum perception, vision foundation model pipelines, and exploratory wireless signal sensing. A RAG component is developed to ground a locally hosted language model in project-specific documentation, including specialized damage level classification criteria to mitigate hallucinations during inference. Controlled comparisons against static few-shot prompting demonstrate that dynamic retrieval improves grounding and factual consistency. We further compare vector-based RAG with a knowledge graph variant constructed via entity-relation extraction, and show that graph-based retrieval produces stronger responses for damage assessment queries requiring cross-document reasoning, motivating hybrid dense, sparse, and graph-aware retrieval. To address limitations of EO imagery under adverse lighting and weather conditions, infrared (IR)/thermal sensing is employed for object detection and segmentation. Our detectors generate candidate detections, yielding improved segmentation of a broad array of objects. Paired IR versus visible spectrum tracking experiments reveal failure modes, motivating multimodal fusion for robust object detection and damage analysis. Vision foundation and vision-language models are leveraged to generate synthetic damage imagery and classify damage severity with high accuracy, supporting training and validation of downstream damage assessment models. Finally, exploratory Wireless-based sensing demonstrates potential to detect presence, motion, and post-event environmental changes where EO and IR sensing are ineffective.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑