发表机构
Beihang University; Beijing Academy of Artificial Intelligence; Peking University; Tsinghua University(北京航空航天大学; 北京人工智能研究院; 北京大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ActiveArena提出模拟器与基准,含35个任务和13种VLA配置,系统评估机器人主动感知,发现记忆采样、容量、本体感觉及子任务监督可提升OOD泛化,为主动感知研究提供统一测试平台。
AI 中文摘要
主动感知和操作对于机器人交互复杂场景至关重要。现有基准难以评估机器人如何以主动方式有效获取并维持记忆中的信息。为此,我们引入了ActiveArena-Sim,一个具有可控视角和大规模工作空间的主动感知模拟器,作为基础。在此基础上,我们提出了ActiveArena-Bench,包含5个细粒度类别下的35个任务,涵盖视觉探索和交互式信息获取。每个任务仅凭被动观察难以解决,需要多轮证据获取和基于记忆的推理。该基准提供了丰富的记忆标注、标准化训练数据以及ID/OOD协议,其中包含不相交场景、未见过的干扰物配置和新颖背景。此外,我们提出了ActiveArena-VLA,一个包含13种视觉-语言-动作配置的模块化套件,用于对主动感知中的记忆写入、记忆容量、本体感觉状态、子任务监督和高层规划进行受控研究。基准结果揭示了显著的ID-OOD差距:在可靠写入策略下,均匀记忆采样、增加记忆容量、本体感觉输入和子任务监督改善了OOD泛化,而规划器引导的记忆管理和决策在使用稀疏记忆时达到了接近最佳配置的性能。因此,ActiveArena为开发和诊断主动感知与操作模型提供了一个统一的测试平台。
英文摘要
Active perception and manipulation are crucial for robots to interact with complex scenes. Existing benchmarks struggle to evaluate how robots effectively acquire and maintain information in memory in an active manner. To this end, we introduce ActiveArena-Sim, an active-perception simulator with controllable viewpoints and large-scale workspaces as the foundation. Built on this, we propose ActiveArena-Bench, which comprises 35 tasks across 5 fine-grained categories, covering visual exploration and interactive information acquisition. Each task is difficult to solve from passive observations alone, requiring multi-round evidence acquisition and memory-based reasoning. The benchmark provides rich memory annotations, standardized training data, and ID/OOD protocols featuring disjoint scenes, unseen distractor configurations, and novel backgrounds. Moreover, we present ActiveArena-VLA, a modular suite of 13 vision-language-action configurations for controlled studies of memory writing, memory capacity, proprioceptive state, subtask supervision, and high-level planning in active perception. Benchmark results reveal a substantial ID-OOD gap: uniform memory sampling, increased memory capacity under reliable write policies, proprioceptive inputs, and subtask supervision improve OOD generalization, while planner-guided memory management and decision-making achieve performance close to the best-performing configuration using only sparse memory. ActiveArena thus provides a unified testbed to develop and diagnose models for active perception and manipulation.
Comments43 pages. Project page: https://leeibo.github.io/ActiveArena