发表机构
Tsinghua University; National University of Singapore(清华大学; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对具身推理现有方法缺乏长程规划的局限,提出ParallelWorld多步测试时缩放框架,通过验证器引导的树搜索实现并行多步轨迹模拟,在ESI-Bench上显著提升了主动感知与推理性能。
AI 中文摘要
具身推理是具身智能的基础能力,是智能体在物理环境中实现自主感知、推理与交互的基础。近期研究已将具身推理范式从静态感知转向动态探索,智能体通过与环境交互获取任务相关信息。然而,现有主动推理方法通常逐步生成探索轨迹,缺乏长程规划;甚至最新的测试时缩放框架也采用近视式单步前瞻,难以应对复杂遮挡空间环境中固有的延迟反馈问题。为解决这一局限,本文提出ParallelWorld,一种用于具身推理的多步测试时缩放框架。与贪婪式单步尝试不同,ParallelWorld使智能体在执行动作前并行模拟并评估多步未来轨迹。具体而言,本文引入验证器引导的树搜索范式:从当前状态出发,ParallelWorld分支为多个并行轨迹并在多步范围内持续展开;在每个模拟步骤中,验证器智能体评估中间状态转移,动态剪除前景不佳的分支并优先选择信息增益最高的路径;多步前瞻模拟完成后,智能体综合长程结果以确定最优动作序列;最后,答案智能体对所选轨迹进行推理以生成最终结果。在ESI-Bench上的大量实验表明,ParallelWorld可持续提升主动感知与推理性能。
英文摘要
Embodied Reasoning constitutes a fundamental capability of embodied intelligence, serving as the basis for autonomous perception, reasoning, and interaction within physical environments. Recent studies have shifted the paradigm of embodied reasoning from static perception toward dynamic exploration, where agents acquire task-relevant information through interactions with the environment. However, existing active reasoning approaches generally generate exploration trajectories incrementally without long-horizon planning. Even recently emerged test-time scaling frameworks often resort to myopic, single-step lookaheads, which struggle to resolve the delayed feedback inherent in complex, occluded spatial environments. To address this limitation, we propose ParallelWorld, a multi-horizon test-time scaling framework for embodied reasoning. Instead of greedy, single-step trials, ParallelWorld empowers agents to simulate and evaluate multi-step future trajectories in parallel before committing to an action. Specifically, we introduce a verifier-guided tree-search paradigm. Starting from the current state, ParallelWorld branches into multiple parallel trajectories and rolls them out continuously across a multi-step horizon. At each simulation step, a verifier agent evaluates the intermediate state transitions, dynamically pruning unpromising branches and prioritizing paths with the highest information gain. Once the multi-step prospective simulation is complete, the agent synthesizes the long-horizon outcomes to commit to the optimal action sequence. Finally, an answer agent performs reasoning over the selected trajectory to produce the final reasoning. Extensive experiments on ESI-Bench demonstrate that ParallelWorld consistently improves active perception and reasoning performance.
CommentsProject Page: https://chen-min-22.github.io/ParallelWorld-page/