AI 中文总结
本研究提出Speech2Grasp框架,以数据高效方式将文本条件抓取检测模型迁移至语音输入,通过轻量级MLP投影器适配,在人形机器人实验中性能优于级联ASR流水线且延迟更低,为文本系统扩展到语音提供实用范式。
AI 中文摘要
人形机器人日益需要多模态理解以实现与人类的自然交互。尽管视觉-语言模型备受关注,但它们通常假设输入是文本而非更自然的语音。本文研究能否以数据高效的方式将成熟的文本条件模型迁移到语音。以ALBEF为例,我们开展诊断分析,结果表明基于轻量级MLP的投影器可有效适配语音,同时保留语义区分性和鲁棒性。基于这些发现,我们提出Speech2Grasp,这是一个用于将文本条件抓取检测高效迁移到语音的框架。真实世界人形机器人实验显示,Speech2Grasp优于级联ASR的流水线,且推理延迟更低。我们的发现为将成熟的文本条件系统扩展到语音提供了实用范式。
英文摘要
Humanoid robots increasingly require multi-modal understanding for natural interaction with humans. Despite the prominence of vision-language models, they generally assume textual rather than the more natural speech inputs. In this paper, we investigate whether a well-established text-conditioned model can be transferred to speech in a data-efficient manner. Using ALBEF as a case study, we conduct diagnostic analyses showing that a lightweight MLP-based projector effectively adapts it to speech, while preserving semantic discrimination and robustness. Motivated by these findings, we introduce Speech2Grasp, a framework for data-efficient transfer of text-conditioned grasp detection to speech. Real-world humanoid robot experiments show that Speech2Grasp outperforms cascaded ASR-based pipeline, while reducing inference latency. Our findings suggest a practical paradigm for extending established text-conditioned systems to speech.