发表机构
Texas A&M University; Purdue University(德克萨斯农工大学; 普渡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对现有视觉-语言-动作模型缺乏未来任务动态推理机制的问题,提出PHR-VLA框架,通过辅助未来头对齐潜在动态,提升了机器人操纵任务的成功率。
AI 中文摘要
视觉-语言-动作模型(VLAs)已展现出通过将语言指令和视觉观测直接映射为动作,实现通用机器人操纵的强大潜力。然而,大多数VLAs主要基于当前观测条件进行动作预测,缺乏对未来任务动态进行推理的显式机制,而这对于精细、接触密集型操纵尤为重要。我们提出PHR-VLA,这一框架通过未来动态的特权潜在表征,使VLAs具备规划视界推理能力。PHR-VLA引入了一个轻量级辅助未来头,在训练期间将VLA的内部表征与从未来观测中提取的潜在动态对齐。评估结果表明,来自腕部相机的局部、以接触为中心的补丁级潜在动态监督,使LIBERO的成功率从84.1%提升至88.4%,现实世界拆卸任务的成功率从63.3%提升至82.5%;来自第三人称相机的补丁级监督也使Meta-World的性能从56.70%提升至57.8%。这些结果表明,特权潜在动态对齐为提升VLA策略的预期推理提供了有效的训练信号。项目网站:https://...
英文摘要
Vision-language-action models (VLAs) have shown strong promise for general-purpose robotic manipulation by mapping language instructions and vision observations directly to actions. However, most VLAs primarily condition action prediction on current observations and lack an explicit mechanism for reasoning over future task dynamics, which is particularly important for fine-grained, contact-rich manipulation. We present PHR-VLA, a framework that enables planning-horizon reasoning in VLAs through privileged latent representations of future dynamics. PHR-VLA introduces a lightweight auxiliary future head that, during training, aligns the VLA's internal representations with latent dynamics extracted from future observations. Evaluation results demonstrate that local, contact-centric, patch-level latent dynamics supervision from the wrist camera improves success rate on LIBERO from 84.1% to 88.4% and on real-world disassembly tasks from 63.3% to 82.5%. Patch-level supervision from a third-person camera also improves performance on Meta-World from 56.70% to 57.8%. These results demonstrate that privileged latent dynamics alignment provides an effective training signal for improving anticipatory reasoning in VLA policies. Project website: \href{https://davoodsz.github.io/PHR-VLA.github.io/}{https://davoodsz.github.io/PHR-VLA.github.io/}