arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.06955cs.ROcs.CV

ROMA:面向真实世界以物体为中心的多感官主动感知的LLM系统

ROMA: LLM System for Real-World Object-Centric Multi-Sensory Active Perception

Ruoxuan Feng, Yutong Chen, Ruihua Song, Huan Yang, Zhongyuan Wang, Guocai Yao, Di Hu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出ROMA,一个基于LLM的主动感知系统,通过推理-互动-反馈循环整合多感官信息,构建大规模数据集ROMI-2K和两阶段训练框架,实现主动获取缺失证据并解决复杂多感官感知任务,为具身智能体奠定感知基础。

中文摘要 AI 辅助

人类天生通过主动过程来理解物理世界。当感官证据不足以推断物理属性时,我们自然地与环境互动,决定缺失哪些信息、如何获取这些信息,以及何时已获得足够的证据。与此形成鲜明对比的是,现有的多感官机器人系统主要整合感官输入,而不是通过互动主动获取缺失的证据。在这项工作中,我们介绍了ROMA,一个基于LLM的系统,用于真实世界以物体为中心的多感官主动感知。ROMA将视觉、音频、触觉和力传感整合到一个推理-互动-反馈循环中。该模型识别缺失的证据并确定目标物体、互动和模态,而物理接口执行选定的互动并收集多感官反馈。为了支持这一能力,我们构建了ROMI-2K,一个大规模的真实世界多感官物体互动数据集,涵盖近2000个物体和6种原子互动,并带有同步的感官反馈。基于这些数据,我们开发了一个两阶段训练框架,对齐感官模态,并使LLM能够评估证据充分性、选择信息丰富的互动,并推理多感官反馈。我们进一步将主动感知表征为感知链,其中获得的证据指导后续互动和推理,并建立了ROMA Bench来评估单属性、长视野多属性和意图驱动的主动感知。实验表明,ROMA能够主动获取缺失的证据,并解决现有方法难以处理的复杂、长链多感官感知任务,为主动多感官具身智能体奠定了坚实的感知基础。

英文摘要

Humans inherently understand the physical world through an active process. When sensory evidence is insufficient to infer physical properties, we naturally interact with the environment by deciding what information is missing, how to acquire it, and when sufficient evidence has been obtained. In stark contrast, existing multi-sensory robot systems mainly integrate sensory inputs rather than actively acquiring missing evidence through interactions. In this work, we introduce ROMA, an LLM-based system for Real-World Object-Centric Multi-Sensory Active Perception. ROMA integrates vision, audio, tactile, and force sensing into a reasoning-interaction-feedback loop. The model identifies missing evidence and determines the target objects, interactions, and modalities, while a physical interface executes the selected interactions and collects the multi-sensory feedback. To support this capability, we construct ROMI-2K, a large-scale real-world multi-sensory object interaction dataset covering nearly 2,000 objects and 6 atomic interactions with synchronized sensory feedback. Building on these data, we develop a two-stage training framework that aligns sensory modalities and equips the LLM to assess evidence sufficiency, select informative interactions, and reason over the multi-sensory feedback. We further characterize active perception as perception chains, where acquired evidence guides subsequent interactions and reasoning, and establish ROMA Bench to evaluate single-attribute, long-horizon multi-attribute, and intent-driven active perception. Experiments show that ROMA can actively acquire missing evidence and solve complex, long-chain multi-sensory perception tasks that existing methods struggle to handle, laying a strong perceptual foundation for active multi-sensory embodied agents.

发表机构

  • Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院)
  • Beijing Key Laboratory of Research on Large Models and Intelligent Governance(北京大模型与智能治理重点实验室)
  • Beijing Academy of Artificial Intelligence(北京人工智能研究院)
  • Beijing Jiaotong University(北京交通大学)
  • State Key Laboratory of Multimedia Information Processing, Peking University(北京大学多媒体信息处理国家重点实验室)
  • AresoX

机构由 AI 辅助整理,请以论文原文为准。

↑