发表机构
Stony Brook University; Salesforce AI Research(石溪大学; Salesforce AI研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SMILE是一种保留架构的接口,通过预测B样条系数生成平滑动作序列,应用于多款VLA模型后,在多个基准及真实实验中提升了长视界VLA执行的准确率与效率
AI 中文摘要
视觉-语言-动作(VLA)模型通过每次调用执行多个动作来降低推理成本,但更长的视界往往会降低准确率,因为原始动作块包含抖动和异常值。我们提出SMILE,这是一种保留架构的接口,可预测B样条系数并将其解码为平滑的动作序列。SMILE仅改变动作表示,能够在保留每个基线的骨干网络和模型规模的同时实现更长的固定视界。我们将SMILE应用于SmolVLA、Evo1、VPP和DAWN,在LIBERO、CALVIN及真实世界实验中提高了准确率和摊销推理效率。SMILE-Evo1在LIBERO上达到98.0%的准确率,同时实现1.1倍加速;SMILE-VPP在CALVIN上达到平均长度4.42,同时实现1.5倍加速。在匹配的10的执行视界下,SMILE-SmolVLA使非边界加速度降低78.6%,速度符号变化率降低42.3%。真实世界xArm测试显示出更高的成功率、更少的掉落和更少的接触。这些结果表明,平滑系数空间生成是实现准确、高效的长视界VLA执行的一条途径。项目页面:this http URL
英文摘要
Vision-Language-Action (VLA) models reduce inference cost by executing multiple actions per call, but longer horizons often degrade accuracy because raw chunks contain jitter and outliers. We introduce SMILE, an architecture-preserving interface that predicts B-spline coefficients and decodes them into smooth action sequences. SMILE changes only the action representation, enabling longer fixed horizons while retaining each baseline's backbone and model scale. We apply SMILE to SmolVLA, Evo1, VPP, and DAWN, improving accuracy and amortized inference efficiency across LIBERO, CALVIN, and real-world experiments. SMILE-Evo1 reaches 98.0% with a 1.1x speedup on LIBERO, while SMILE-VPP reaches an average length of 4.42 with a 1.5x speedup on CALVIN. At a matched execution horizon of 10, SMILE-SmolVLA reduces non-boundary acceleration by 78.6% and velocity sign-change rate by 42.3%. Real-world xArm tests show higher success, fewer drops, and fewer contacts. These results establish smooth coefficient-space generation as a route to accurate, efficient long-horizon VLA execution. Project page: jongwoopark7978.github.io/smilevla
CommentsSubmitted to IEEE Robotics and Automation Letters (RA-L)