arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13605cs.AIcs.RO

用于具身消歧的主动感知

Active Perception for Embodied Disambiguation

Yiwei Liu, Luwei Yang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对具身环境中目标歧义问题,提出以主动观测为核心的主动感知框架,结合视觉语言模型实现信息获取与用户意图澄清,经真实机器人实验验证有效。

中文摘要 AI 辅助

自然语言为机器人提供了灵活的任务接口,但具身环境中的目标歧义不仅源于用户意图,还可能是当前观测中缺少与任务相关的物理证据导致的。现有交互式消歧方法主要通过询问用户获取额外信息,而遮挡、受限视角、不可读文本和未观测目标则要求机器人主动改变其观测方式。我们提出了一种用于具身目标消歧的主动感知框架,该框架以主动观测作为信息获取的核心,使用视觉语言模型基于累积的视觉证据和交互信息,决定是继续观测、请求澄清还是完成目标选择。主动观测既可以直接恢复缺失的判别性证据,也可以揭示对象名称、标签和语义属性,从而在仍需用户澄清时优化其效果。真实机器人实验表明,该框架在统一的具身消歧流程中结合了物理信息获取和用户意图澄清。

英文摘要

Natural language provides robots with a flexible task interface, but target ambiguity in embodied environments arises not only from user intent; it can also result from missing taskrelevant physical evidence in the current observation. Existing interactive disambiguation methods primarily obtain additional information by asking the user, whereas occlusion, restricted viewpoints, unreadable text, and unobserved targets require the robot to actively change its observation. We propose an active-perception framework for embodied target disambiguation that uses active observation as the backbone for information acquisition and uses a vision-language model to decide, on the basis of accumulated visual evidence and interaction information, whether to continue observing, request clarification, or complete target selection. Active observation can both directly recover missing discriminative evidence and reveal object names, labels, and semantic attributes, thereby improving user clarification when it remains necessary. Real-robot experiments show that the framework combines physical information acquisition and userintent clarification within a unified embodied disambiguation process.

发表机构

  • School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)理工学院)
  • Shenzhen Research Institute of Big Data(深圳大数据研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑