arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MulRobBench:用于安全且符合安全策略的多模态无人机智能体的决策级基准测试

MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents

Belal S. Alsinglawi, Weizheng Wang, Junyi Wu, Lianhai Lin, Merouane Debbah, Izzat Alsmadi

arXiv 2607.23870首次发表:更新:

发表机构

Zayed University; The University of Adelaide; University of Emergency Management; Khalifa University; Texas A&M University–San Antonio(扎耶德大学; 阿德莱德大学; 应急管理大学; 哈利法大学; 德州农工大学圣安东尼奥分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

智慧城市中无人机需在复杂条件下决策,MulRobBench作为决策级基准测试,将多模态观测、安全策略等集成,含多任务样本和评分维度,评估结合多种方式,给出模型得分,通过研究找出决策不稳定原因,为无人机多模态决策提供可重复基准。

AI 中文摘要

智慧城市空域正将无人机从被动传感平台转变为必须在观测退化和语言模糊情况下遵循操作规则的网络物理决策者。现有无人机和多模态基准测试评估感知、导航、协作和推理,但很少评估关键决策过程中物理证据、协议约束和行动风险是否保持耦合。我们引入MulRobBench,这是一个用于智慧城市环境中视觉-语言-行动(VLA)无人机智能体的离线、协议条件基准测试。它将真实无人机多模态观测、协议级安全策略和行动级网络物理安全集成到一个统一评估框架中。该基准测试包含3024个样本,跨越17个任务分类节点和四个阶段的12个评分维度。评估结合语义评分和结构诊断。研究发现最佳语义协议决策得分仅为0.5141,最佳严格平均评分维度准确率为0.1599。模态消融研究证实视觉和文本输入都会影响决策,分析确定了决策不稳定的主要原因。MulRobBench为现实操作约束下可信的多模态无人机决策提供了可重复的基准测试。

英文摘要

In IoT-enabled smart-city settings, Uncrewed Aerial Vehicles (UAVs) are evolving from passive sensing platforms into cyber-physical decision makers that must respect operational rules under degraded observations and ambiguous language. Existing UAV and multimodal benchmarks cover aerial perception, navigation, collaboration, and task reasoning, but rarely test whether physical evidence, protocol constraints, and action risk stay coupled at critical decisions. We introduce MulRobBench, an offline, protocol-conditioned benchmark for Vision-Language-Action (VLA) UAV agents that links real UAV multimodal observations, protocol-level security-policy constraints, and action-level cyber-physical safety within an auditable decision contract. The evaluation set contains 3,024 samples spanning 17 task-taxonomy nodes and 12 metric scoring dimensions, organized around context understanding, multimodal evidence arbitration, degradation-aware reasoning, and risk-aware action planning. MulRobBench reports controlled semantic scores alongside strict structural diagnostics for policy compliance, formatting, unsafe actions, parsing, and dimension-level validity. Across 17 uniformly audited models, the best semantic protocol-decision score reaches 0.5141 and the best strict mean scoring-dimension accuracy reaches 0.1599. A matched 20-anchor modality-removal study changes 4-15 action selections per model, showing both visual and textual inputs influence decisions while the strongest input condition varies across metrics. Per-dimension and conditional analyses identify modality-trust selection, constraint extraction, strong glare, missing data, and high-entropy operator shorthand as principal sources of action instability. The central challenge is thus stable coupling of degraded evidence, security-policy constraints, and risk-bearing action, not isolated scene recognition.

Comments22 pages, 18 figures, 17 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑