arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AdaHVLA:面向长时程视觉-语言-动作执行的适应性框架

AdaHVLA: Adaptive Harnesses for Long-Horizon Vision-Language-Action Execution

Junyi Tang, Jie Peng, Zezhen Ding, Yuan Shen, Tianlong Chen

arXiv 2609.29204首次发表:更新:

发表机构

UNC; USTC; HKUST; CUHK(北卡罗来纳大学; 中国科学技术大学; 香港科技大学; 香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AdaHVLA通过解耦多智能体适应过程,利用机器人经验优化协调策略,将长时程VLA任务成功率显著提升,并支持跨任务持续适应。

AI 中文摘要

视觉-语言-动作(VLA)模型在局部控制和指令跟随方面表现出色,但在需要持久记忆和规划的长时程任务中往往力不从心。任务框架通过保留任务历史并跟踪执行各阶段的进展,为智能体推理提供持久上下文。为融合这些互补能力,我们提出AdaHVLA,一种适应性框架,通过机器人经验优化基于代码的协调策略,使智能体的推理和记忆与VLA执行更好地对齐。其解耦的多智能体适应过程将证据分析、框架修订和行为评估分离到不同的工作上下文中,利用可测试的协调假设指导修订,并通过随后的试运行评估其预期效果。一个带状态的修订图将执行证据、假设、修订和观察到的效果联系起来,保留替代框架和适应记忆,以指导跨重复尝试的改进以及跨任务和环境的持续适应。在仿真中,AdaHVLA将NaVILA-LH上的平均测试成功率从22.5%提升至高达57.5%,并在三个VLA骨干网络上将操作测试成功率较初始框架提升最多30.8个百分点。真实世界部署进一步展示了适应后的策略如何支持跨任务阶段的稳定执行。

英文摘要

Vision-language-action (VLA) models offer strong local control and instruction following but often struggle with long-horizon tasks requiring persistent memory and planning. Task harnesses provide persistent context for agent reasoning by retaining task history and tracking progress across execution stages. To bring these complementary capabilities together, we introduce AdaHVLA, an adaptive harness that refines code-based coordination policies through robot experience to better align agent reasoning and memory with VLA execution. Its decoupled multiagent adaptation process separates evidence analysis, harness revision, and behavioral assessment into distinct working contexts, using testable coordination hypotheses to guide revisions and subsequent rollouts to assess their predicted effects. A stateful revision graph links execution evidence, hypotheses, revisions, and observed effects, preserving alternative harnesses and adaptation memory to guide refinement across repeated attempts and continued adaptation across tasks and environments. In simulation, AdaHVLA raises mean test success on NaVILA-LH from 22.5\% to as high as 57.5\% and improves manipulation test success across three VLA backbones by up to 30.8 percentage points over the initial harness. Real-world deployment further illustrates how the adapted policies support stable execution across task stages.

Comments9 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑