发表机构
Robotics Institute, Carnegie Mellon University; The Hong Kong University of Science and Technology (Guangzhou)(卡内基梅隆大学机器人研究所; 香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对固定视点下机器人操作信息缺失问题,提出ActiveScale框架,通过模型、数据、硬件协同设计,实现VLA模型主动感知,提升任务成功率。
AI 中文摘要
当固定视点导致任务相关信息被遮挡或不可见时,主动感知对于机器人操作至关重要。然而,使视觉-语言-动作(VLA)模型能够在变化的视点间进行推理并主动获取信息丰富的观测仍然具有挑战性。我们提出了ActiveScale,一个通过协调模型、数据和硬件设计来推进主动感知的框架。我们的模型为VLA增加了历史视频观测和显式相机位姿监督,使用逐帧位姿标记和轻量级预测头来关联不同视点间的观测,并支持对场景的一致理解。为了从人类活动中自然存在的相机运动中学习,我们引入了一种可扩展的人-机器人中期训练方案,使用1000小时的自我中心和机器人数据,使模型适应时间输入和位姿监督。我们进一步介绍了主动感知移动操作平台(AMP),一个通过单操作员远程操作支持主动感知和移动操作的机器人平台,能够可扩展地收集协调视点变化和操作的演示。实验表明,在主动感知任务上成功率有所提高,而消融研究验证了相机位姿感知建模和自我中心中期训练的贡献。这些组件共同为研究和开发机器人操作中的主动感知提供了集成基础。
英文摘要
Active perception is essential for robotic manipulation when fixed viewpoints leave task-relevant information occluded or unobserved. However, enabling vision-language-action (VLA) models to reason across changing viewpoints and actively acquire informative observations remains challenging. We present ActiveScale, a framework that advances active perception through coordinated model, data, and hardware designs. Our model augments a VLA with historical video observations and explicit camera-pose supervision, using per-frame pose tokens and a lightweight prediction head to associate observations across viewpoints and support a coherent understanding of the scene. To learn from the camera motion naturally present in human activity, we introduce a scalable human--robot mid-training recipe using 1000 hours of egocentric and robotic data, adapting the model to temporal inputs and pose supervision. We further introduce Active-perception Mobile-manipulation Platform (AMP), a robotic platform that supports active perception and mobile manipulation through single-operator teleoperation, enabling scalable collection of demonstrations that coordinate viewpoint changes and manipulation. Experiments demonstrate improved success rates on active-perception tasks, while ablation studies validate the contributions of camera-pose-aware modeling and egocentric mid-training. Together, these components provide an integrated foundation for studying and developing active perception in robotic manipulation.
Commentsactive-scale.github.io