AI 中文总结
该研究针对智能家居智能体的提示注入安全问题,构建了现实场景基准测试PromptShield-Home,对比了三类防御层的性能,发现多智能体调解的上限性能最优,提出应采用学习路由与传感器融合方案保障安全。
AI 中文摘要
智能家居助手越来越多地使用可直接感知视频和音频的多模态大语言模型(MLLMs),这引发了智能家居特有的安全问题:智能体能否将真实用户命令与环境或外部来源的内容、电视语音、屏幕文本或偷听的对话区分开,这些内容仅看起来像命令?我们推出了PromptShield-Home,这是一个涵盖受话人歧义、屏幕/音频注入、健康监测器误触发、混合居住以及合法命令下限的现实智能家居场景试点基准,并用它来比较三个抽象层:传统检测器(L0)、单个MLLM智能体(L1;视觉、视觉+自动语音识别(ASR)、视听)和多智能体调解(L2;投票、角色专家、跨模型仲裁)。由于标签分布偏向弃权(不执行),总体准确率具有误导性,始终阻止预测器的得分为82%,因此我们分别报告不安全执行率和安全完成率。两种范式的失败方式相反:检测器对所有内容都采取行动,而每种MLLM配置都过度拒绝,几乎不完成任何真实命令,且在所有情况下都遗漏了真正的跌倒事件。关键的是,它们的正确集合互不重叠:始终选择正确层的神谕达到94.1%,而最佳单层的得分为76.5%。我们将此报告为上限,而非系统——未实现任何路由——并认为智能家居智能体安全最好通过学习路由和传感器融合来实现,而非用MLLM取代检测器。
英文摘要
Smart-home assistants increasingly use multimodal large language models (MLLMs) that perceive video and audio directly. This raises a safety question specific to the home: can the agent tell a genuine user command from ambient or externally-sourced content, television speech, on-screen text, or an overheard conversation, that merely looks like a command? We introduce PromptShield-Home, a pilot benchmark of realistic smart-home scenarios spanning addressee ambiguity, screen/audio injection, health-monitor false triggers, mixed occupancy, and a legitimate-command floor, and use it to compare three abstraction layers: traditional detectors (L0), a single MLLM agent (L1; vision, vision+ASR, and audio-visual), and multi-agent mediation (L2; voting, role specialists, cross-model arbitration). Because the label distribution is skewed toward inaction, aggregate accuracy is misleading, a constant always-block predictor scores 82%, so we report unsafe-execution and safe-completion rates separately. The two paradigms fail in opposite ways: detectors act on everything, while every MLLM configuration over-refuses, completing almost no genuine command and missing a true fall in every case. Crucially, their correct sets are disjoint: an oracle that always picks the right layer reaches 94.1%, against 76.5% for the best single layer. We report this as an upper bound, not a system - no router is implemented - and argue that home-agent safety is best served by learned routing and sensor fusion, not by replacing detectors with an MLLM.
CommentsThis work has been accepted as a poster to UbiComp 2026