离线多模态大语言模型用于空中作战决策支持
Offline Multimodal Large Language Models for Decision Support in Air Operations
浏览论文内容
中文总结 AI 辅助
本研究探索离线多模态大语言模型作为空中作战决策支持工具,提出检索增强架构,试点显示其性能与人类相当且用时更短,为AI辅助评估建立基线。
中文摘要 AI 辅助
空中作战依赖于复杂规则、既定程序以及在有限连接和严格安全约束下的时间关键分析。在此类环境中,分析人员必须将书面条令与图像相结合,且往往无法访问外部计算资源。本文研究离线大语言模型作为决策支持工具,部署在隔离和受限环境中,通过自然语言交互使分析人员能够获取可追溯至原始来源的条令知识。我们描述了一种适用于无互联网连接运行的模块化检索增强架构,支持来自技术手册的文本和图像输入。作为评估该架构的第一步,我们报告了一项针对巴西空军四名图像分析人员的试点研究,结合了(i)基于其电子目标识别条令的条令知识评估,在同一测试中比较人类与所提系统的表现,以及(ii)对无AI辅助情况下手动生成侦察目标报告(Reconnaissance Mission Report - REMIR)所涉及的认知工作负荷的测量。结果表明手动任务要求高,尤其在脑力需求(6.0/7)和努力程度(5.0/7)方面,而所提系统与人类得分(8/10)持平,并在7.1分钟内完成评估(而人类平均为26.5分钟),为未来AI辅助评估建立了基线。最后,我们描述了一个未来评估协议,以系统比较手动与AI辅助工作流程。
英文摘要
Air operations rely on complex rules, established procedures, and time-critical analysis under limited connectivity and strict security constraints. In such environments, analysts must combine written doctrine with images, often without access to external computing resources. This paper studies offline large language models as decision support tools, deployed in isolated and restricted environments to give analysts access to doctrinal knowledge that remains traceable to its original sources through natural language interaction. We describe a modular retrieval-augmented architecture suitable for operation without Internet connectivity, supporting both text and image input from technical manuals. As a first step toward evaluating this architecture, we report a pilot study with four image analysts of the Brazilian Air Force, combining (i) a doctrinal knowledge assessment based on their electronic-target identification doctrine, comparing human and proposed system performance on the same test, and (ii) a measurement of the cognitive workload involved in manually producing a reconnaissance target report (Relatório de Missão de Reconhecimento - REMIR) without AI assistance. The results show a demanding manual task, especially in terms of mental demand (6.0/7) and effort (5.0/7), while the proposed system matches the human score (8/10) and completes the assessment in 7.1 minutes (compared to a human average of 26.5 minutes), establishing a baseline for future AI-assisted evaluation. Finally, we describe a future evaluation protocol to systematically compare manual and AI-assisted workflows.
发表机构
- Instituto de Estudos Avançados(高级研究所)
机构由 AI 辅助整理,请以论文原文为准。