PonderPounce:作为机器人控制回合上下文引擎的预训练多模态大语言模型(MLLM)
PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control
浏览论文内容
中文总结 AI 辅助
PonderPounce复用预训练MLLM的因果上下文作为机器人记忆,由System2 MLLM(Ponder)和System1 VLA(Pounce)组成,在RoboMME、RoboCasa-DC等数据集上的性能优于现有基线,实现了低延迟的机器人控制。
中文摘要 AI 辅助
多模态大语言模型(MLLM)可整合长视觉历史、在部分可观测条件下推理并从少量示例中推断行为。然而视觉-语言-动作(VLA)模型通常继承预训练表示,未将这种上下文能力用作回合记忆。依赖记忆的策略通过专用历史机制解决这一缺口,而PonderPounce则复用MLLM的原生因果上下文作为机器人记忆。Ponder是System2 MLLM,在其原生因果上下文中积累回合观测、演示和先验认知,可生成子目标文本和内部使用的演示推理;Pounce是System1 VLA,直接接收当前观测、指令和本体感觉,通过Ponder-Pounce接口异步仅接收最新的连续认知标记及其年龄。二者端到端联合训练,无专用记忆模块或单独的桥接预训练。优化后的服务使认知刷新的p50延迟为78ms,动作模型调用的p50延迟为25ms,支持20Hz动作回放。在RoboMME上,使用基础规模训练数据时,在相同的Pounce架构和接口下,9B规模的PonderPounce准确率达60.83%,0.8B规模的达50.04%,而FrameSamp+Modul为44.51%,仅用当前观测的π₀.₅为17.93%;使用9倍数据时,其准确率达75.54%,而FrameSamp+Modul为57.88%。在RoboCasa-DC上,相同接口仅从动作监督中学习,准确率达12.5%,而最强的已发表演示条件基线为11.6%,当认知被替换为学习到的空状态时,准确率降至8.6%。
英文摘要
Multimodal large language models (MLLMs) can integrate long visual histories and infer behavior from a few examples, yet vision-language-action models rarely use this capacity as episode memory. Instead of a purpose-built memory module, PONDERPOUNCE reuses an MLLM's native causal context. PONDER, a pretrained System 2 MLLM, integrates episode history and demonstrations to produce continuous cognition. POUNCE, a System 1 action model, asynchronously conditions control on the newest cognition and its age. Both are jointly trained end to end without separate bridge pretraining. Optimized per-call inference on an H100 achieves p50 latencies of 78 ms for cognition-only refresh and 25 ms for action-model invocation. On RoboMME, PONDERPOUNCE achieves 60.83% success at the base data scale and 75.54% with 9x data, compared with 44.51% and 57.88% for FrameSamp+Modul. At base scale, scaling PONDER from 0.8B to 9B adds 6.71 percentage points with the POUNCE architecture unchanged. A separately trained 9B PONDER without execution history achieves only 26.21% under matched supervision. PONDERPOUNCE also achieves 12.5% success on RoboCasa-DC and demonstrates real-world applicability on four tasks under asynchronous execution, with 60.98% mean success versus 40.67% for FrameSamp+Modul.
发表机构
- Seoul National University(首尔大学)
机构由 AI 辅助整理,请以论文原文为准。