arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29489cs.CRcs.CV

SpatialTrust:安全认证中环境风险识别的基准测试

SpatialTrust: A Benchmark for Environmental Risk Recognition in Secure Authentication

  • University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

Junbin Lu, Hsiang-Wei Huang, Saesha Wadhwa, Yu Ting Hsu, Jenq-Neng Hwang

AI总结:

本文提出SpatialTrust基准测试评估MLLM在安全认证场景中环境风险识别等五项能力,发现现有模型表现有限,还引入SpatialTrustGuard流程提升了Qwen3-VL-30B-A3B-Instruct的性能。

AI中文摘要:

视觉环境风险识别在安全认证中发挥重要作用,用户周围环境可能泄露敏感信息或带来潜在安全风险。然而,现有多模态大语言模型(MLLM)评估很少检验模型能否在空间接地的认证场景中可靠识别、定位并解释此类风险。本文提出SpatialTrust,这是一个用于评估安全认证中环境风险识别的问答基准测试,它评估五项互补能力:敏感因素检测、直接因素识别、间接因素识别、直接因素解释和间接因素解释。我们评估了专有及开源MLLM,发现当前模型性能有限,尤其在理解和解释间接风险方面,表明空间风险感知仍是MLLM的一项具有挑战性的能力。此外,我们引入SpatialTrustGuard,一种结构化问答与审计流程,将Qwen3-VL-30B-A3B-Instruct的整体性能从36.78%提升至41.12%。研究结果强调,需专用基准测试和结构化推理方法提升MLLM在安全认证中的可信度。

英文摘要:

Visual environmental risk recognition plays an important role in secure authentication, where a user's surroundings may reveal sensitive information or introduce potential security risks. However, existing evaluations of multimodal large language models (MLLMs) rarely examine whether models can reliably recognize, localize, and explain such risks in spatially grounded authentication scenarios. We present SpatialTrust, a question-answering benchmark for evaluating environmental risk recognition in secure authentication. SpatialTrust assesses five complementary abilities: sensitive factor detection, direct factor identification, indirect factor identification, direct factor explanation, and indirect factor explanation. We evaluate both proprietary and open-source MLLMs and find that current models show limited performance, especially in understanding and explaining indirect risks, indicating that spatial risk awareness remains a challenging capability for MLLMs. In addition, we introduce SpatialTrustGuard, a structured QA-and-audit pipeline that improves Qwen3-VL-30B-A3B-Instruct from 36.78% to 41.12% overall. Our findings highlight the need for dedicated benchmarks and structured inference methods to improve the trustworthiness of MLLMs in secure authentication.

补充信息

↑