arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GuardianBench:面向具身智能中潜在上下文风险的同场景指令对比基准

GuardianBench: A Same-Scene Instruction-Contrastive Benchmark for Latent Contextual Risk in Embodied AI

Zhesheng Zhang, Jiahao Lu, Wei Liu, Cong Pan, Jianhua Yang, Yixiang Chen, Hongyuan Yu, Mengqi Zhang, Kailin Lyu, Zhumin Chen, Keji He

arXiv 2608.21928首次发表:更新:

发表机构

Shandong University; National University of Singapore; Nanjing University of Aeronautics and Astronautics; Institute of Automation, Chinese Academy of Sciences; Xiaomi Corporation(山东大学; 新加坡国立大学; 南京航空航天大学; 中国科学院自动化研究所; 小米公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出基于国际安全标准的同场景指令对比基准GuardianBench,发现VLMs对指令不敏感,用轻量级目标VLOS可提升其安全推理性能。

AI 中文摘要

在具身智能领域,安全风险可能是潜在的:一条看似无害的指令和一个安全的场景,只有当二者结合时才会变得危险。现有研究通过改变视觉上下文或评估执行时的动态特性来推进具身安全,但固定场景、仅改变指令这一互补维度的探索仍不充分。我们提出GuardianBench,这是一个基于国际安全标准的指令对比基准,通过3024个指令-场景示例(按不同危险类别组织为同场景的安全/不安全对比对)来分离这种潜在上下文风险。对最先进的视觉语言模型(VLMs)进行基准测试后发现,模型存在对指令不敏感的判定问题:在给定场景下,模型会不成比例地批准两条指令;在主要模型中,平均对准确率仅为24.1%。我们开展了系统性的原理审计,定位出主要故障:模型无法关联区分安全与不安全组合的指令相关线索。作为训练后案例研究,Verdict Log-Odds Supervision(VLOS,一种轻量级的判定级目标函数)显著提升了开放权重骨干模型的性能。综上,我们提出的潜在上下文风险任务范式、基于标准的对比基准构建、对级和原理级故障诊断,以及基准支持的判定校准,使GuardianBench成为一套用于暴露和改进潜在上下文风险下指令-场景组合安全推理的可控评估套件。

英文摘要

In embodied AI, safety risk can be latent: a benign instruction and a safe scene become hazardous only when composed. Prior work has advanced embodied safety by varying visual contexts or evaluating execution-time dynamics, but the complementary axis of fixing the scene and varying only the instruction remains underexplored. We introduce GuardianBench, an instruction-contrastive benchmark grounded in international safety standards that isolates this latent contextual risk through 3,024 instruction-scene examples organized as same-scene Safe/Unsafe contrastive pairs across various hazard categories. Benchmarking state-of-the-art vision-language models (VLMs) reveals instruction-insensitive verdicts: models disproportionately approve both instructions under a given scene; across the primary models, average pair accuracy is only 24.1%. Our systematic rationale audit localizes the dominant failure: models fail to bind the instruction-relevant cues that differentiate safe from unsafe compositions. As a post-training case study, Verdict Log-Odds Supervision (VLOS), a lightweight verdict-level objective, substantially improves performance on open-weight backbones. Together, our latent contextual risk task formulation, standards-grounded contrastive benchmark construction, pair-level and rationale-level failure diagnosis, and benchmark-enabled verdict calibration establish GuardianBench as a controlled evaluation suite for exposing and improving safety reasoning over instruction-scene compositions under latent contextual risk.

Comments21 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑