arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18392cs.RO

DistAL:基于距离的优势学习用于VLA微调

DistAL: Distance-based Advantage Learning for VLA Fine-Tuning

  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

Reece O'Mahoney, Ioannis Havoutis

AI总结:

DistAL利用嵌入空间距离作为奖励,改进VLA微调中的价值函数,提升下游任务成功率,并在仿真和真实硬件上验证。

AI中文摘要:

视觉-语言-动作模型(VLAs)近年来通过将大语言模型(LLMs)的语义理解与流匹配策略的精确控制相结合,彻底改变了机器人操作领域。优势条件化是一种近期技术,通过在部署数据上训练价值函数,并利用该函数训练优势条件化策略,从而迭代改进VLA。以往的工作仅应用简单、低信息的成功/失败奖励,这使得价值函数无法区分不同质量的状态,只能根据任务进展程度进行判断。受分布外(OOD)检测方法探索的启发,我们引入了基于距离的优势学习(DistAL),该方法通过使用嵌入空间距离作为奖励,产生更具信息量的价值函数,进而提高下游任务成功率。我们在系列仿真基准测试和真实硬件上的灵巧双臂操作任务中验证了我们的方法。

英文摘要:

Vision-language-action models (VLAs) have trans- formed the field of robotic manipulation in recent years by combining the semantic understanding of LLMs with the precise control of flow-matching policies. Advantage conditioning is a recent technique that iteratively improves VLAs by training a value function on deployment data and using this to train an advantage-conditioned policy. Previous works have only applied simple, low-information success/failure rewards, which leave the value function unable to distinguish states of differing quality beyond how far along the task they appear. Motivated by an exploration of out-of-distribution (OOD) detection methods, we introduce Distance-based Advantage Learning (DistAL), which, by using an embedding space distance as a reward, produces a more informative value function and subsequently a higher downstream task success rate. We validate our method on a series of simulation benchmarks and dexterous bi-manual manipulation tasks on real hardware.

↑