面向具身智能体的任务状态视界的编译与基准测试
Compiling and Benchmarking Task-State Horizons for Embodied Agents
浏览论文内容
中文总结 AI 辅助
该研究定义任务状态视界(TSH),推出RoboGraph编译器并发布含588个情节的基准,发现多数先进智能体模型在高TSH下维护、更新任务相关状态存在差距。
中文摘要 AI 辅助
前沿智能体模型正越来越多地被用作长视界具身任务的高层规划器。现有的机器人基准测试已在长视界评估方面取得进展,但主要通过动作序列长度和子任务复杂度来表征难度,却忽略了一个独特挑战:智能体必须跟踪由自身探索和环境动态所引发的不断变化的任务相关世界状态。我们将智能体必须跟踪的任务相关状态转换的范围定义为任务状态视界(task-state horizon,TSH)。为评估智能体性能如何随TSH变化,我们推出RoboGraph,一种机器人任务编译器,它能将状态转换依赖关系转换为可执行的符号图。具体而言,RoboGraph从空间和时间因果依赖关系中构建任务状态视界,包括任务执行期间意外故障和干预所引发的依赖关系。基于RoboGraph,我们发布了一个包含84个场景中588个情节的基准测试,这些情节具有不同的TSH。在语义和视觉闭环环境中对15种先进智能体模型的实验表明,大多数模型在高要求的TSH下表现不佳,揭示了其在长视界下维护、探索和更新任务相关状态方面存在巨大差距。
英文摘要
Frontier agentic models are increasingly deployed as high-level planners for long-horizon embodied tasks. Existing robotic benchmarks have advanced long-horizon evaluation, but primarily characterize difficulty through action-sequence length and subtask complexity, overlooking a distinct challenge: agents must track evolving task-relevant world states induced by both their exploration and environmental dynamics. We define the span of task-relevant state transitions that an agent must track as task-state horizon (TSH). To evaluate how agent performance varies with TSH, we introduce RoboGraph, a robotic task compiler that translates state-transition dependencies into executable symbolic graphs. Specifically, RoboGraph constructs task-state horizons from spatial and temporal causal dependencies, including those induced by unexpected failures and interventions during task execution. Building on RoboGraph, we release a benchmark comprising 588 episodes across 84 scenes with varying TSHs. Experiments evaluating 15 advanced agentic models in both semantic and visual closed-loop environments show that most models struggle with demanding TSHs, revealing substantial gaps in maintaining, exploring, and updating task-relevant state over long horizon.