arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

连接前沿推理与机器人执行:从自主演示生成到密集语言监督

Bridging Frontier Reasoning and Robot Execution: From Autonomous Demonstration Generation to Dense Language Supervision

Bosung Kim, Alexander Trevithick, Ruiyi Wang, Prithviraj Ammanabrolu

arXiv 2610.03615首次发表:更新:

发表机构

University of California, San Diego; NVIDIA(加州大学圣迭戈分校; 英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出自主演示生成与密集语言监督两种互补方法,连接前沿推理与低延迟机器人执行,在长时域任务中提升指令遵循性能,并验证其协同有效性。

AI 中文摘要

近期前沿模型的进展使得机器人仅需少量演示即可进行操控,但高推理延迟限制了其在实时机器人控制中的应用。为弥合这一差距,我们研究了两种互补的方法,将前沿推理与低延迟本地执行相连接。首先,我们使用前沿模型自主生成演示,以补充人类演示,用于训练快速的本地策略。我们通过加入纠正性演示片段来增强其上下文示例,这些片段展示了如何从物理错误中恢复,从而提高了生成的可靠性。随着成功示例在上下文中积累,生成时间和成本下降,这为更高效的数据收集提供了一条路径。在部署时,一个框架将前沿模型生成的指令与快速本地策略相结合,在保持前沿模型指导和纠正动作能力的同时,实现高效执行。由于低延迟执行被委托给本地策略,瓶颈转移到其可靠遵循前沿模型多样化指令的能力上。我们的第二座桥梁引入了跨三个嵌套粒度(原始、原子和复合)的密集语言监督,每个层级都包含多方面描述。在RoboCasa 365和BEHAVIOR-1K中的长时域任务中,指令随执行进展而变化,结合监督在oracle和前沿模型指令下均取得了最高性能,展示了通过策略语言接口更可靠的指令遵循。最后,我们在一个结合语义规划和操作并在固定时间预算内的填字游戏任务中评估了这两座桥梁。这些结果支持自主演示生成和密集语言监督作为连接前沿推理到低延迟本地执行的互补组件。

英文摘要

Recent advances in frontier models enable robot manipulation from only a few demonstrations, but high inference latency limits their use for real-time robot control. To bridge this gap, we study two complementary approaches that connect frontier reasoning with low-latency local execution. First, we use a frontier model to autonomously generate demonstrations that supplement human demonstrations for training a fast local policy. We augment its in-context examples with corrective demonstration segments that show how to recover from physical errors, improving generation reliability. Generation time and cost decrease as successful examples accumulate in context, suggesting a path toward more efficient data collection. At deployment, a harness combines frontier-generated instructions with a fast local policy, enabling efficient execution while preserving the frontier model's ability to guide and correct actions. With low-latency execution delegated to the local policy, the bottleneck shifts to its capacity to reliably follow the frontier model's diverse instructions. Our second bridge introduces dense language supervision across three nested granularities---primitive, atomic, and composite---with multi-aspect descriptions at each level. Across long-horizon tasks in RoboCasa 365 and BEHAVIOR-1K, where instructions change as execution progresses, the combined supervision achieves the highest performance under both oracle and frontier-model instructors, demonstrating more reliable instruction following through the policy's language interface. Finally, we evaluate both bridges together on a crossword task that combines semantic planning and manipulation within a fixed time budget. These results support autonomous demonstration generation and dense language supervision as complementary components for connecting frontier reasoning to low-latency local execution.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑