arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32336cs.ROcs.HC

WSM-Aware HRI:一种物联网增强的框架,用于在LLM指导下早期检测和规范引导修复故障

WSM-Aware HRI: An IoT-Enhanced Framework for Early Detection and Norm-Guided Repair of Failures with LLM Guidance

Hanlin Zhang, Yuquan Wang, Tianwei Zhang, Zhenglong Sun

首次发表
浏览论文内容

中文总结 AI 辅助

提出WSM-Aware HRI框架,利用物联网和LLM将人机交互故障统一为世界状态不匹配,实现意图形成阶段的主动检测与规范引导修复,在十个案例中验证了有效性。

中文摘要 AI 辅助

人机交互(HRI)中的故障仍然是机器人在现实环境中部署的主要障碍。以往的工作通常将故障视为孤立的技术缺陷,或侧重于事后恢复行为。实际上,许多故障的产生是因为人类和机器人在当前世界状态的假设上存在不一致。我们提出了WSM-Aware HRI,一种物联网增强的模块化框架,将各种HRI故障统一为人类指令隐含假设与机器人基于多模态感知和数字增强构建的接地世界模型之间的世界状态不匹配(WSMs)。利用大型语言模型(LLM)将隐含假设显式化,将其映射到少量不匹配类型,并指定验证所需证据以对照机器人世界状态进行核查。WSM-Aware HRI将故障处理从执行时恢复转变为意图形成期间的主动不匹配检测,从而能够根据安全性、规范合规性和多用户协调进行干预,并提供透明的解释。我们在十个涵盖视觉和潜在状态不匹配的日常案例中评估了不匹配识别。系统能够准确产生预期输出结果,消融实验表明,可靠的识别依赖于适当的接地表示和面向验证的细化。这些结果表明,将交互故障视为显式的世界状态不匹配能够更早地检测即将发生的故障,并为将外部证据和社会约束整合到人机交互中提供了一种原则性机制。

英文摘要

Human-robot interaction (HRI) failures remain a major barrier to deploying robots in real-world environments. Prior work often treats failures as isolated technical faults or focuses on post-hoc recovery behaviors. In practice, many breakdowns arise because humans and robots operate under inconsistent assumptions about the current world state. We propose WSM-Aware HRI, an IoT-enhanced modular framework that unifies diverse HRI breakdowns as World-State Mismatches (WSMs) between a human's instruction-implied assumptions and a robot's grounded world model built from multimodal perception and digital augmentation. A Large Language Model (LLM) is used to make implicit assumptions explicit, map them to a small set of mismatch types, and specify the evidence needed for verification against the robot's world state. WSM-Aware HRI shifts failure handling from execution-time recovery to proactive mismatch detection during intention formation, enabling interventions guided by safety, norm compliance, and multi-user coordination with transparent explanations. We evaluate mismatch identification in ten everyday cases spanning both visual and latent-state mismatches. The system can accurately produce the expected output results, and ablations show that reliable identification depends on appropriate grounding representations and verification-oriented refinement. These results indicate that treating interaction breakdowns as explicit world-state mismatches enables earlier detection of impending failures and offers a principled mechanism for integrating external evidence and social constraints into human-robot interaction.

↑