arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RoboInter1.5:用于具身世界建模和机器人操作的整体中间表示套件

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation

Ziqin Wang, Hao Li, Weijun Wang, Junhao Cai, Jia Zeng, Yilun Chen, Jiangmiao Pang, Si Liu

arXiv 2607.18709首次发表:更新:

AI 中文总结

研究针对现有机器人数据集问题,基于RoboInter1.0提出RoboInter1.5套件,含多种中间表示资源及相关任务,通过广泛评估证明其为中间表示提供统一时空框架,将其作为双向接口,规范动作空间并约束物理模拟器潜在展开。

AI 中文摘要

现有机器人数据集难以构建,特定于实体,且缺乏用于可推广推理、执行或长期环境动态模拟所需的细粒度结构注释。基于之前的RoboInter1.0,提出RoboInter1.5,这是一个用于机器人操作和具身世界建模的扩展且整体的中间表示套件。它提供了以密集的面向操作的中间表示为中心的数据、基准和模型统一资源。具体包括含多种注释的RoboInter-Data,引入相关任务的RoboInter-VQA,研究表示对动作执行益处的RoboInter-VLA,以及利用中间表示预测未来世界状态的RoboInter-World。广泛评估表明RoboInter1.5为中间表示提供了统一的时空框架,将其概念化为双向接口。

英文摘要

Existing robot datasets remain expensive to curate, embodiment-specific, and insufficiently annotated with the fine-grained structure required for generalizable reasoning, execution, or long-horizon environment dynamics simulation. Building on our prior work, RoboInter1.0, we present RoboInter1.5, an extended and holistic suite of intermediate representations for both robotic manipulation and embodied world modeling. RoboInter1.5 provides a unified resource of data, benchmarks, and models centered on dense manipulation-oriented intermediate representations. Specifically, RoboInter-Data contains over 230k manipulation episodes across 571 scenes with dense per-frame annotations covering more than ten types of intermediate representations, including subtasks, primitive skills, object and gripper grounding, segmentation, affordance, grasp poses, contact points, motion traces, etc. Built upon these annotations, RoboInter-VQA introduces spatial and temporal embodied VQA tasks to benchmark and improve the intermediate-representation reasoning capabilities of our RoboInter-VLM. RoboInter-VLA further studies how such representations benefit action execution through implicit, explicit, and modular plan-then-execute paradigms. To better model the physical world, we further introduce RoboInter-World, which leverages intermediate representations as structured conditioning signals for controllable prediction of future world states. Extensive evaluations demonstrate that RoboInter1.5 provides a unified spatiotemporal scaffolding for intermediate representations. Rather than treating intermediate representations merely as interpretable signals, RoboInter1.5 conceptualizes them as a bidirectional interface that both regularizes low-level action spaces and constrains the latent rollouts of open-world physical simulators.

Comments28 pages. arXiv admin note: substantial text overlap with arXiv:2602.09973

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑