Token Expand-Merge: Training-Free Token Compression for Vision-Language-Action Models
令牌扩展-合并:面向视觉-语言-动作模型的免训练 令牌压缩
机构 * College of Science and Engineering, Hamad Bin Khalifa University(1 科学与工程学院,哈马德·本·卡西姆大学) ; Mohamed bin Zayed University of Artificial Intelligence(2 摩萨·本·扎耶德人工智能大学) ; College of Computer Science and Technology, Zhejiang University(3 计算机科学与技术学院,浙江大学)
AI总结 TEAM-VLA通过动态令牌扩展与合并机制,实现无需训练的视觉-语言-动作模型高效推理,提升速度并保持任务性能。
Comments 8 pages, 5 figures