发表机构
Fraunhofer IPA(弗劳恩霍夫应用研究促进协会IPA研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出首个面向人机协作中LLM编排器的安全基准,基于MCP架构和ISO标准定义安全不变量与合规分类,通过文本、模拟和物理三层评估发现模型家族决定安全底线,上下文管理可减少行为问题但可能增加违规。
AI 中文摘要
大型语言模型(LLMs)越来越多地被用于通过自然语言接口编排机器人行为,然而目前尚无基准来评估其在人机协作中作为安全感知决策者的可靠性。与强制执行二元允许/拒绝决策的确定性安全系统不同,基于LLM的编排器表现出一个合规谱系,从过度合规(拒绝安全动作)到完全违反安全。本文首次提出了人形机器人协作中LLM编排器的安全基准环境,该环境基于模型上下文协议(MCP)架构构建,安全不变量基于ISO 10218-2:2025防护措施。该基准定义了五个可测试的安全不变量、一个四级合规分类法(正确合规、过度合规、欠合规、完全违反)以及一个三层评估流程(文本提示、模拟传感器-执行器回路、在Unitree G1 EDU人形机器人上的物理验证)。我们报告了第一层结果:三个云后端(Claude Haiku 4.5、GPT-4o-mini、Gemini 2.5 Flash)和一个本地开源权重基线(qwen3:8b),在完整上下文和滑动窗口预算条件下进行了40次100轮会话,而模拟和物理层仍在进行中。我们发现:(1)模型家族决定了安全底线,Claude和Gemini保持零或接近零违规,而GPT-4o-mini每会话最多违规13次;(2)上下文管理将两个失败轴分离,使每个云后端的平均行为问题减少42-57%,同时使GPT-4o-mini的违规次数几乎翻倍(从每会话3.8次增至7.2次);(3)比例合规(将移动速度限制在规则指定的最大值而不是拒绝)仅一致出现在Gemini中;初步模拟层重现了模型排名和GPT-4o-mini失败模式的反转。
英文摘要
Large Language Models (LLMs) are increasingly employed to orchestrate robot behavior through natural-language interfaces, yet no benchmark exists to evaluate their reliability as safety-aware decision makers in human-humanoid collaboration. Unlike deterministic safety systems that enforce binary allow/deny decisions, LLM-based orchestrators exhibit a compliance spectrum ranging from overcompliance (refusing safe actions) to full safety violations. This paper introduces the first safety benchmarking environment for LLM orchestrators in human-humanoid collaboration, built on a Model Context Protocol (MCP)-based architecture with safety invariants grounded in ISO 10218-2:2025 protective measures. The benchmark defines five testable safety invariants, a four-level compliance taxonomy (correct compliance, overcompliance, undercompliance, full violation), and a three-layer evaluation pipeline (text prompting, simulated sensor-actuator loops, and physical validation on a Unitree G1 EDU humanoid). We report Layer-1 results: three cloud backends (Claude Haiku 4.5, GPT-4o-mini, Gemini 2.5 Flash) and a local open-weights baseline (qwen3:8b) across 40 100-turn sessions under full-context and sliding-window budget conditions, while the simulation and physical layers remain ongoing. We find that (1) model family determines the safety floor, as Claude and Gemini remain at or near zero violations while GPT-4o-mini commits up to 13 per session, (2) context management dissociates two failure axes, reducing mean behavioral issues by 42-57% for every cloud backend while nearly doubling GPT-4o-mini's violations (3.8 to 7.2 per session), and (3) proportional compliance, clamping movement speed to the rule-specified maximum rather than refusing, emerges consistently only in Gemini; the preliminary simulation layer reproduces the model ranking and the GPT-4o-mini failure-mode inversion.
Comments8 pages, CBS 2026