arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Attacca:面向长时程具身智能体的状态连续性目标导向控制

Attacca: Goal-Directed Control under State Continuity for Long-Horizon Embodied Agents

Gyusik Seo, Jaehong Yoon

arXiv 2610.07785首次发表:更新:

发表机构

Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长时程具身任务中目标不可见导致策略失效的问题,提出Attacca方法,通过上下文解耦目标采样、目标掩码预测和行为阶段条件化,在Minecraft任务中将成功率提升至基线1.7-7倍。

AI 中文摘要

具身智能体的一个核心能力是通过一系列相互依赖的任务来完成复杂目标。然而,现有的视觉目标条件策略通常是在目标已可见的孤立交互场景中进行评估,因此无法捕捉连续长时程任务执行过程中出现的条件。在此类设置中,每个任务都从前一个任务留下的状态开始:智能体可能结束于不同的位置和朝向,世界可能已被修改,下一个交互目标可能位于当前视野之外。因此,依赖此类策略的智能体在无法将目标锚定于当前观察时,可能难以继续执行下一个任务。为应对这一挑战,我们提出了Attacca,一种新的方法,它使用与执行环境解耦的目标图像,在完整的搜索到交互轨迹上训练视觉目标条件策略。Attacca采用上下文解耦目标采样,将每个演示与来自另一个世界的类别兼容掩码目标图像配对,消除了直接的场景和姿态对应关系。它通过目标掩码预测头学习稠密的当前视图锚定,提供超越动作模仿的辅助监督。我们进一步引入了行为阶段条件化,教导策略区分搜索、接近和交互阶段,并根据执行进展调整其控制。我们在Minecraft中的多个短时程和长时程具身任务上评估了Attacca。我们的方法实现了39.0%-47.5%的干净成功率,比最强基线提升了1.7-2.4倍。在长时程任务上,它达到了54%、30%和28%的完成率,实现了高达7倍的改进。

英文摘要

A central capability of embodied agents is to accomplish complex objectives through sequences of interdependent tasks. Yet existing visual goal-conditioned policies underlying these agents are typically evaluated on isolated interactions where the target is already visible, and thus do not capture the conditions that arise during continuous long-horizon task execution. In such settings, each task begins from the state left by the previous one: the agent may end at a different position and orientation, the world may have been modified, and the next interaction target may lie outside the current field of view. As a result, agents relying on such policies may struggle to proceed to the next task when they cannot ground their target in the current observation. To address this challenge, we propose Attacca, a new approach that trains visual goal-conditioned policies on complete search-to-interact trajectories using goal images decoupled from the execution environment. Attacca uses context-decoupled goal sampling to pair each demonstration with a class-compatible masked goal image from another world, removing direct scene and pose correspondence. It learns dense current-view grounding through a target-mask prediction head, providing auxiliary supervision beyond action imitation. We further introduce behavioral-phase conditioning that teaches the policy to distinguish Search, Approach, and Interact stages and adapt its control as execution progresses. We evaluate Attacca on multiple short- and long-horizon embodied tasks in Minecraft. Our method achieves 39.0-47.5% clean success, improving over the strongest baseline by 1.7-2.4x. On long-horizon tasks, it attains 54%, 30%, and 28% completion, yielding up to a 7x improvement.

CommentsProject page: https://attacca-project.github.io

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑