ExoActor:作为可泛化的交互人形控制的外向视频生成
ExoActor: Exocentric Video Generation as Generalizable Interactive Humanoid Control
浏览论文内容
中文总结 AI 辅助
本文提出ExoActor框架,利用大规模视频生成模型的泛化能力,通过第三人称视频生成统一接口建模交互动态,将任务指令和场景上下文转化为可执行的人形行为序列,展示了其在新场景中的泛化能力。
中文摘要 AI 辅助
人形控制系统近年来取得了显著进展,但建模机器人、其周围环境和任务相关物体之间流畅的交互丰富行为仍是一个基本挑战。这种困难源于需要在大规模层面联合捕捉空间上下文、时间动态、机器人动作和任务意图,这与传统监督不匹配。我们提出ExoActor,一种新的框架,利用大规模视频生成模型的泛化能力来解决这个问题。ExoActor的关键思想是使用第三人称视频生成作为统一接口来建模交互动态。给定一个任务指令和场景上下文,ExoActor合成出合理的执行过程,这些过程隐含编码了机器人、环境和物体之间的协调互动。此类视频输出然后通过一个管道转换为可执行的人形行为,该管道估计人类运动并通过通用运动控制器执行,产生一个任务条件的行为序列。为了验证所提出的框架,我们将其实现为端到端系统,并展示了其在没有额外现实数据收集的情况下对新场景的泛化能力。此外,我们通过讨论当前实现的局限性和未来研究的有希望方向,说明了ExoActor如何提供一种可扩展的方法来建模交互丰富的机器人行为,可能为生成模型进一步推进通用机器人智能开辟新途径。
英文摘要
Humanoid control systems have made significant progress in recent years, yet modeling fluent interaction-rich behavior between a robot, its surrounding environment, and task-relevant objects remains a fundamental challenge. This difficulty arises from the need to jointly capture spatial context, temporal dynamics, robot actions, and task intent at scale, which is a poor match to conventional supervision. We propose ExoActor, a novel framework that leverages the generalization capabilities of large-scale video generation models to address this problem. The key insight in ExoActor is to use third-person video generation as a unified interface for modeling interaction dynamics. Given a task instruction and scene context, ExoActor synthesizes plausible execution processes that implicitly encode coordinated interactions between robot, environment, and objects. Such video output is then transformed into executable humanoid behaviors through a pipeline that estimates human motion and executes it via a general motion controller, yielding a task-conditioned behavior sequence. To validate the proposed framework, we implement it as an end-to-end system and demonstrate its generalization to new scenarios without additional real-world data collection. Furthermore, we conclude by discussing limitations of the current implementation and outlining promising directions for future research, illustrating how ExoActor provides a scalable approach to modeling interaction-rich humanoid behaviors, potentially opening a new avenue for generative models to advance general-purpose humanoid intelligence.
发表机构
- Beijing Academy of Artificial Intelligence (BAAI)(北京人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。