arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.13605cs.RO

关于VLA微调的阶段信息接口的实证研究

An Empirical Study on Stage-Information Interfaces for VLA Fine-Tuning

Yingwei Ji

首次发表
浏览论文内容

中文总结 AI 辅助

研究长期操作中VLA微调的阶段信息接口,通过分段动作注释等方法,在直接微调与延续微调下和GR00T N1.6比较,发现不同接口表示和训练安排下阶段信息对策略成功率影响不同。

中文摘要 AI 辅助

在长期操作中,一个高级指令可涵盖多个动作阶段。我们使用分段动作注释作为全任务指令和VLA动作块之间的中间表示。一个进度模块跟踪活动阶段,而动作策略接收阶段信息,其形式可以是当前阶段文本或机器人状态中的归一化序数阶段索引。我们在直接微调以及从全任务指令基线进行的延续微调下,将这些接口与LIBERO - 10上的GR00T N1.6进行比较。直接微调时,全任务指令、当前阶段文本和序数阶段状态的平均成功率分别为57.45%、50.24%和54.36%,表明明确的阶段信息不会自动改善策略。延续微调时,相应均值为49.07%、50.00%和53.75%,序数阶段状态在所有三次配对运行中均超过其他两者。观察到的益处因接口表示和训练安排而异。

英文摘要

One high-level instruction in long-horizon manipulation can cover several action stages. We use segmented action annotations as an intermediate representation between the full-task instruction and VLA action chunks. A progress module tracks the active stage, while the action policy receives stage information either as current-stage text or as a normalized ordinal stage index in robot state. We compare these interfaces with GR00T N1.6 on LIBERO-10 under direct fine-tuning and continuation fine-tuning from a full-task instruction baseline. Under direct fine-tuning, full-task instruction, current-stage text, and Ordinal Stage-State achieve mean success rates of 57.45%, 50.24%, and 54.36%, respectively, showing that explicit stage information does not automatically improve the policy. Under continuation, the corresponding means are 49.07%, 50.00%, and 53.75%, with Ordinal Stage-State exceeding both alternatives in all three paired runs. The observed benefit differs across interface representations and training arrangements.

补充信息

↑