发表机构
Tsinghua University; Tongji University(清华大学; 同济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ProactiveVLA通过主动环境探索获取可复用知识,在部署时增强具身记忆,在LIBERO-Pro和RoboCasa365上优于基线,并在LIBERO-Pro Goal-T上以48%完成率大幅超越19%的最先进基线。
AI 中文摘要
快速适应新环境要求机器人从有限的体验中获取关于局部对象、状态和交互的有用知识。将推理智能体与冻结的视觉-语言-动作模型(VLA)相结合的系统可以通过执行反馈和记忆进行适应,这使得体验的选择对其有效性至关重要。对目标任务的重复练习可能完善一个熟悉的解决方案,同时使与变化条件相关的其他交互未经测试。我们提出了ProactiveVLA,它利用主动环境探索来获取可复用的知识,以用于部署时的适应。在完成初始任务后,智能体将剩余的交互预算分配给自提议的目标,这些目标涵盖对象可供性、状态改变交互以及交互的组合。它验证执行结果,并将任务导向和探索性的体验整合到指导后续规划和控制的记忆中。在相同的交互预算下,ProactiveVLA在LIBERO-Pro和RoboCasa365 Composite-Seen上优于基线方法。在LIBERO-Pro Goal-T上,评估期间最多允许一次VLA原语调用,ProactiveVLA完成了48%的实例,而最先进的任务细化基线仅为19%。
英文摘要
Rapid adaptation to a new environment requires a robot to acquire useful knowledge about local objects, states, and interactions from limited experience. Systems that combine a reasoning agent with a frozen vision-language-action model (VLA) can adapt through execution feedback and memory, making the choice of experience central to their effectiveness. Repeated practice of a target task may refine a familiar solution while leaving other interactions relevant to changed conditions untested. We introduce ProactiveVLA, which uses proactive environment exploration to acquire reusable knowledge for deployment-time adaptation. After completing an initial task, the agent allocates the remaining interaction budget to self-proposed goals covering object affordances, state-changing interactions, and compositions of interactions. It verifies execution outcomes and consolidates both task-directed and exploratory experience into memory that guides subsequent planning and control. ProactiveVLA outperforms the baselines under the same turn budget on LIBERO-Pro and RoboCasa365 Composite-Seen. On LIBERO-Pro Goal-T, with at most one VLA primitive invocation allowed during evaluation, ProactiveVLA completes 48% of instances, compared with 19% for the state-of-the-art task-refinement baseline.