发表机构
Northwestern; UCSB; UIUC; CMU(西北大学; 加州大学圣塔芭芭拉分校; 伊利诺伊大学厄巴纳-香槟分校; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长任务中具身智能体进度估计依赖上下文的问题,构建ContextProgress-Bench基准并验证上下文关键性,提出ProgressCompass智能体循环,利用通用VLM提供上下文,使冻结PRM误差降低63%、排名一致性提高76%。
AI 中文摘要
具身智能体现在承担越来越长的任务。对于长任务,仅知道任务最终成功或失败并不能说明太多问题;过程中的每一步都很重要。进度奖励模型(PRMs)在每一步对任务已完成的进度进行评分,并作为密集奖励、验证器和监控器。然而,在长任务中,仅凭当前帧往往无法判断任务已完成的进度,因为进度取决于之前发生的情况。我们将此问题称为上下文相关的进度估计。现有的进度估计基准主要集中于短任务,其进度可以从当前观察中读取,而PRMs在需要上下文时能否估计进度仍未得到充分探索。因此,我们构建了ContextProgress-Bench,包含24个操作任务,共120个片段。该基准涵盖三种设置:(i)状态回忆(State Recall),其中所需信息在之前出现但不在当前帧中;(ii)序列跟踪(Sequence Tracking),其中步骤遵循固定顺序,因此进度需要知道哪些步骤已完成以及下一步是什么;(iii)重复消歧(Recurrence Disambiguation),其中外观相似的帧位于非常不同的进度阶段。然后我们进行配对诊断:每个PRM在两次运行中保持相同的输入格式,其中一次运行的指令整合了正确的上下文。即使读取整个历史的PRMs在估计进度时也会迷失方向,但在正确上下文下,同样的五个模型将其进度误差降低了77-82%。因此,具身PRMs并非不能估计进度,而是缺乏正确上下文时会迷失方向。因此,我们提出ProgressCompass,一个自主智能体循环,重新定向现有的PRM,并使用当前通用的视觉语言模型(VLMs)来提供PRM所需的上下文。在该循环中,相同的冻结PRM将其进度误差降低了63%,并将其排名一致性提高了76%。有了这样的指南针,PRMs在更长、更复杂的任务上能更好地估计进度。
英文摘要
Embodied agents now take on ever longer tasks. For long tasks, knowing only whether a task finally succeeds or fails says little; the steps along the way matter. Progress Reward Models (PRMs) score how far a task has come at every step, and serve as dense rewards, verifiers and monitors. Yet in long tasks the current frame alone often cannot tell how far the task has come, because progress depends on what happened before. We call this problem context-dependent progress estimation. Existing benchmarks on progress estimation mostly focus on short tasks whose progress can be read from the current observation, and whether PRMs can estimate progress when context is needed remains underexplored. We therefore build ContextProgress-Bench, with 24 manipulation tasks for 120 episodes. The benchmark covers three settings: (i) State Recall, where information needed for progress appeared earlier but is not in the current frame; (ii) Sequence Tracking, where steps follow a fixed order, so progress requires knowing which steps are done and which comes next; and (iii) Recurrence Disambiguation, where look-alike frames sit at very different progress. We then run a paired diagnosis: each PRM keeps the same input format in both runs, and in one run its instruction integrates the right context. Even PRMs that read the entire history get lost in estimating progress, yet with the right context the same five models cut their progress error by 77-82%. Embodied PRMs are thus not incapable of progress estimation, but lost without the right context. We therefore propose ProgressCompass, an autonomous agentic loop that reorients an existing PRM and uses current general-purpose VLMs to supply the context the PRM needs. Wrapped in the loop, the same frozen PRM cuts its progress error by 63% and raises its rank agreement by 76%. With such a compass, PRMs estimate progress far better on longer, more complex tasks.
CommentsProject page: https://andyzworks.github.io/progresscompass/