OmniSmartHome:智能家居智能体的多模态推理基准
OmniSmartHome: A Multimodal Reasoning Benchmark for Smart-Home Agents
浏览论文内容
中文总结 AI 辅助
OmniSmartHome是一个多模态智能家居基准,将语音请求与视觉和空间音频情境配对,评估16个全模态大语言模型,发现仅靠语音时表现强,需多模态推理时性能下降,并提供PROME基线以提升性能。
中文摘要 AI 辅助
智能家居助手预期能够处理日常生活中出现的多样化、现实请求。在此类交互中,用户通常依赖周围的多模态情境——指向物体或提及所见所闻,从而使得仅凭语言表达的请求变得不明确。然而,现有的智能家居基准仅通过语言表达用户请求,导致依赖情境的现实请求未被充分探索。为弥补这一差距,我们引入了OmniSmartHome,一个多模态智能家居基准,其中每个语音请求均与周围的视觉和空间音频情境配对,提供互补线索以消解不明确的请求。OmniSmartHome包含1,360个合成情境和272个真实世界情境。我们评估了16个全模态大语言模型(Omni-LLMs),并揭示,当仅凭语音足以传达用户意图时,它们表现强劲,但当解析意图需要推理多模态情境线索时,性能显著下降。作为简单的智能体基线,我们提供了PROME(用于多模态证据收集的程序性记忆),该基线为智能体配备专门的视听感知工具和程序性记忆以编排其使用。PROME在六个Omni-LLMs上普遍提升了性能。演示和示例可在以下网址获取:此https URL
英文摘要
Smart-home assistants are expected to handle diverse, realistic requests that arise in daily life. In such interactions, users often rely on the surrounding multimodal context-pointing at objects or referring to what they see or hear, leaving their requests underspecified in language alone. Existing smart-home benchmarks, however, express user requests solely through language, leaving context-dependent real-world requests underexplored. To bridge this gap, we introduce OmniSmartHome, a multimodal smart-home benchmark where each spoken request is paired with the surrounding visual and spatial-audio context, providing complementary cues to disambiguate underspecified requests. OmniSmartHome comprises 1,360 synthetic and 272 real-world episodes. We evaluate 16 omnimodal large language models (Omni-LLMs) and reveal that, while they perform strongly when speech alone sufficiently conveys the user's intent, performance drops substantially when resolving it requires reasoning over multimodal contextual cues. As a simple agent baseline, we provide PROME (PROcedural Memory for multimodal Evidence gathering), which equips agents with specialized audio-visual perception tools and procedural memory for orchestrating their use. PROME generally improves performance across six Omni-LLMs. Demos and examples are available at https://omni-smart-home.github.io