arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25487cs.AIcs.CV

CoTinyVLA:用于十亿参数以下视觉-语言-动作模型的思维链蒸馏

CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model

Minhyeok Lee, Chiyoung Kim, Chanhoe Gu, Seongrok Kim, Sanghyuk Roy Choi, Donghwan Hwang, Donghun Ryu, Seokhyun Kim

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对VLA模型内存需求大问题,提出CoTinyVLA模型。通过双视图时间输入、分层思维链蒸馏、释义增强三个组件,在LIBERO-Plus基准测试中取得优异成绩,证明结构化监督可让小参数主干超越其他模型。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型将自然语言命令转换为机器人动作序列,但LIBERO-Plus鲁棒性基准测试中的领先系统使用30亿到70亿参数的主干,其内存需求可能超过嵌入式机器人预算。我们提出了CoTinyVLA,这是一个基于Qwen3.5-0.8B主干的0.9B参数动作模型,通过构建监督而非扩大模型来获得鲁棒性。三个组件针对问题的不同方面:每步16个历史帧的双视图时间输入,带有文本相机和时间标记;从35B教师模型进行分层思维链(CoT)蒸馏,得到任务阶段、夹爪状态和下一个子动作的情节级计划和块级思维跨度;以及释义增强,将40个基本命令扩展为800个变体。在LIBERO-Plus上,跨越七个扰动维度的10030个扰动任务,CoTinyVLA在空间上达到90.8%,在对象上达到87.3%,在目标上达到86.6%,在长任务上达到80.7%,在所有四个套件上领先最强的7B基线4.7、2.8、15.9和3.0分,每个差距区间都不包括零。增益集中在基准测试最难的方面:在已发布的11个基线中,没有一个在任何套件的机器人初始状态上超过53.2%,而CoTinyVLA在目标上达到73.6%,最强基线为39.9%。消融实验表明,这三个组件可按扰动轴分离,在匹配的图像预算下,帧在两个相机之间以及跨时间的划分本身就占8.6分。闭环推理在分配的GPU内存为2.25 GiB时达到峰值,配对干预表明情节计划是有承载作用的:用空跨度或矛盾跨度替换它会使成功率损失40到45分。因此,结构化监督使0.9B主干超过了所有其他模型。

英文摘要

Vision-Language-Action (VLA) models translate natural-language commands into robot action sequences, but leading systems on the LIBERO-Plus robustness benchmark use three- to seven-billion-parameter backbones whose memory demands can exceed embedded robotic budgets. We present CoTinyVLA, a 0.9B-parameter action model on a Qwen3.5-0.8B backbone that obtains that robustness by structuring supervision instead of enlarging the model. Three components target different axes of the problem: dual-view temporal input of 16 history frames per step with textual camera and time markers; hierarchical chain-of-thought (CoT) distillation from a 35B teacher into an episode-level Plan and a chunk-level Think span over task phase, gripper state and next subaction; and paraphrase augmentation expanding 40 base commands into 800 variants. On LIBERO-Plus, spanning 10,030 perturbed tasks across seven perturbation dimensions, CoTinyVLA reaches 90.8% on Spatial, 87.3% on Object, 86.6% on Goal and 80.7% on Long, leading the strongest 7B baseline on all four suites by 4.7, 2.8, 15.9 and 3.0 points, with every margin interval excluding zero. The gains concentrate on the hardest axes of the benchmark: across the eleven published baselines none exceeds 53.2% on Robot Initial States in any suite, whereas CoTinyVLA reaches 73.6% on Goal against 39.9% for the strongest baseline. Ablations show the three components to be separable by perturbation axis, and at a matched image budget how frames are divided between the two cameras and across time accounts for 8.6 points on its own. Closed-loop inference peaks at 2.25 GiB of allocated GPU memory, and paired interventions show the episode Plan to be load-bearing: replacing it with an empty or contradictory span costs 40 to 45 points of success. Structured supervision thus lets a 0.9B backbone exceed all of them. Code: https://github.com/BrainJellyPie/CoTinyVLA

发表机构

  • Chung-Ang University(韩国中央大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑