arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19630cs.AIcs.CLcs.RO

从意图到行动:车辆语音指令授权中的LLM安全性基准测试

From Intent to Action: Benchmarking LLM Safety in Vehicle Voice Command Authorization

Diba Afroze, Xingli Zhang, Yazhou Tu, Xiali Hei

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出一个202场景基准,评估LLM在车辆语音指令授权中的安全性,发现结构化决策虽提升一致性,但无法消除错误执行,需独立执行层保障安全。

中文摘要 AI 辅助

大型语言模型(LLMs)正日益集成到车辆语音助手中。但将自然语言请求与车辆功能相连接,会产生一个安全关键的授权问题。在执行指令之前,系统必须选择是执行、拒绝、澄清、要求确认、转交手动控制、触发紧急响应,还是不进行工具调用。据我们所知,先前的评估并未在说话者角色、认证状态、车辆状态和工具可用性等方面孤立这一行动前决策。我们引入了一个包含202个场景的基准测试,并采用七类分类法下的参考决策。我们使用决策一致性和安全特定错误指标评估了两个本地开放权重模型和三个基于API的LLM。一致性范围从Llama 3.2 3B的40.1%到Gemini 3.1 Pro Preview的89.1%。基于API的模型得分在83.2%到89.1%之间,它们之间没有统计学上的显著差异。即使是这些模型,在161个非执行场景中也会产生两到三次错误执行,并且在确认和手动控制决策中仍存在持续错误。在结构化授权策略下,受控的Llama 3.2 3B消融实验将一致性提高到40.1%,而仅使用模式和通用安全基线的结果为28.2-29.2%,但这并未消除错误执行。因此,结构化的LLM决策不足以作为独立的安全机制,部署需要一个独立的执行层,在调用任何车辆功能之前验证工具权限和车辆状态约束。

英文摘要

Large language models (LLMs) are increasingly integrated into vehicle voice assistants. But linking natural-language requests to vehicle functions creates a safety-critical authorization problem. Before executing a command, the system must choose whether to execute, refuse, clarify, require confirmation, defer to manual control, trigger an emergency response, or make no tool call. To our knowledge, prior evaluations do not isolate this pre-action decision across speaker role, authentication status, vehicle state, and tool availability. We introduce a 202-scenario benchmark with Reference Decisions under a seven-class taxonomy. We evaluate two local open-weight models and three API-based LLMs using Decision Alignment and safety-specific error metrics. Alignment ranges from 40.1% for Llama 3.2 3B to 89.1% for Gemini 3.1 Pro Preview. The API-based models score between 83.2% and 89.1%, with no statistically significant differences among them. Even these models produce two to three False Executes among 161 non-execution scenarios, and persistent errors remain in confirmation and manual-control decisions. A controlled Llama 3.2 3B ablation increases alignment to 40.1% under the structured authorization policy, versus 28.2-29.2% under schema-only and generic-safety baselines, but it does not eliminate False Executes. Structured LLM decisions are therefore insufficient as a standalone safety mechanism, and deployment requires an independent enforcement layer that verifies tool permissions and vehicle-state constraints before invoking any vehicle function.

发表机构

  • University of Louisiana at Lafayette(路易斯安那大学拉法叶分校)
  • Auburn University(奥本大学)

机构由 AI 辅助整理,请以论文原文为准。

↑