视觉-语言-动作模型的连续条件化:增强肌电与视觉任务描述符
Continuous Conditioning of VLAs with Augmenting EMG and Visual Task Descriptors
浏览论文内容
中文总结 AI 辅助
本文提出EC-VLA和VA-VLA两种模型,分别利用肌电信号和视觉注释作为补充任务条件,在杂乱场景中显著提升VLA模型性能,证明超越语言的任务条件化具有潜力。
中文摘要 AI 辅助
视觉-语言-动作(VLA)模型尽管具有多模态输入,却严重依赖语言来描述任务信息。我们假设状态空间中的其他模态可能为补充任务条件化提供机会,这在杂乱或模糊场景中可能尤其相关。为验证这一假设,我们引入了两个调优模型:(1)电生理条件化VLA(EC-VLA),将8通道肌电包络作为连续条件输入,与本体感觉向量拼接;(2)视觉注释VLA(VA-VLA),将视觉分割注释纳入图像输入。在跨三名参与者评估的立方体选择任务中,EC-VLA在无杂乱、分布内条件下与语言提示基线匹配,并在杂乱、分布外场景中显著优于基线。类似地,VA-VLA在分布内场景中较语言提示基线有适度改进,在杂乱、分布外试验中有显著提升。这些结果共同为超越语言的任务条件化的潜在益处提供了有力证据。
英文摘要
Vision-Language-Action (VLA) models rely strongly on language for describing task information, despite having multimodal inputs. We hypothesize that other modalities in the state space may present opportunities for supplemental task conditioning, which may be particularly relevant in cluttered or otherwise ambiguous scenes. We introduce two tuned models to test this hypothesis: (1) an electrophysiology-conditioned VLA (EC-VLA) that incorporates 8-channel electromyography envelopes as continuous conditioning input concatenated to the proprioceptive vector, and (2) a visually-annotated VLA (VA-VLA) that incorporates visual segmentation annotations to the image inputs. On a cube-selection task evaluated across three participants, EC-VLA matches a language-prompted baseline in uncluttered, in-distribution conditions and substantially outperforms it in cluttered, out-of-distribution scenes. Similarly, VA-VLA shows modest improvements over a language-prompted baseline in in-distribution scenes with substantial improvement in cluttered, out-of-distribution trials. Together, these results provide strong evidence for the potential benefit of task-conditioning beyond language.
发表机构
- Johns Hopkins Applied Physics Laboratory(约翰斯·霍普金斯大学应用物理实验室)
机构由 AI 辅助整理,请以论文原文为准。