BinoGen:面向具身视觉感知与学习的自我中心双目数据规模化生成
BinoGen: Scaling egocentric binocular data for embodied visual perception and learning
浏览论文内容
中文总结 AI 辅助
提出BinoGen框架,自动生成大规模具身感知的自我中心双目视觉数据,含超2000万标注图像,提升真实感知任务性能,并支持研究具身形态对感知学习的影响。
中文摘要 AI 辅助
具身视觉感知依赖于通过与环境的持续互动所积累的时间上连贯的视觉体验。然而,收集大规模自我中心双目观测数据并附带密集标注仍然成本高昂且困难重重。此外,视觉体验不仅受环境影响,还受观察者具身形态的影响,包括观察高度、视场角、双目几何结构以及场景中的运动方式。为应对这些挑战,我们提出了BinoGen,一个用于在室内环境中自动生成大规模、具身感知的自我中心双目视觉体验的框架。BinoGen通过生成式场景合成、概率化物体实例化、外观随机化、随机轨迹生成以及可配置的双目相机设置,联合建模环境与观察者的变化。该框架生成同步的双目视频,并附带密集的多模态监督信息,包括深度图、光流、表面法向量、语义图、物体坐标和相机位姿。利用BinoGen,我们构建了一个包含超过2000万张带标注图像的数据集,用于监督学习。我们展示了BinoGen的两种互补用途。首先,将BinoGen数据纳入训练能够持续提升真实世界视觉感知性能,包括深度估计、目标检测和视频目标跟踪。其次,来自相同环境的成对人类启发和鼠类启发观测数据,使得能够受控地研究观察者具身形态如何影响感知学习。具身特定的适应显著提升了性能,而联合训练则使单一模型能够在两种具身形态下均表现良好。综合这些结果表明,大规模、可控的视觉体验能够提升具身感知能力。
英文摘要
Embodied visual perception relies on temporally coherent visual experience accumulated through continuous engagement with the environment. However, collecting large-scale egocentric binocular observations together with dense annotations remains costly and difficult. Moreover, visual experience is shaped not only by the environment but also by the embodiment of the observer, including viewing height, field of view, binocular geometry, and motion through the scene. To address these challenges, we present BinoGen, an automated framework for generating large-scale, embodiment-aware egocentric binocular visual experiences in indoor environments. BinoGen jointly models environmental and observer variation through generative scene synthesis, probabilistic object instantiation, appearance randomization, stochastic trajectory generation, and configurable binocular camera setups. The framework produces synchronized binocular videos together with dense multimodal supervision, including depth maps, optical flow, surface normals, semantic maps, object coordinates, and camera poses. Using BinoGen, we construct a dataset comprising more than 20 million annotated images for supervised learning. We demonstrate two complementary utilities of BinoGen. First, incorporating BinoGen data consistently improves real-world visual perception, including depth estimation, object detection, and video object tracking. Second, paired human-inspired and mouse-inspired observations from the same environments enable controlled investigation of how observer embodiment affects perceptual learning. Embodiment-specific adaptation substantially improves performance, while joint training enables a single model to perform competitively across both embodiments. Together, these results demonstrate that large-scale, controllable visual experience can improve embodied perception...
发表机构
- College of Biological Sciences, China Agricultural University(中国农业大学生物学院)
- Beijing Institute for Brain Research, Chinese Academy of Medical Sciences & Peking Union Medical College(中国医学科学院北京协和医学院北京脑研究所)
- Chinese Institute for Brain Research, Beijing (CIBR)(北京脑科学与类脑研究中心)
机构由 AI 辅助整理,请以论文原文为准。