SocialVLA:面向VLA操作中基于人类反应失败检测与恢复的社会感知网关
SocialVLA: A Social Perception Gateway for Human-Reaction-Based Failure Detection and Recovery in VLA Manipulation
- Skolkovo Institute of Science and Technology(斯科尔科沃科学技术研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
SocialVLA通过融合人类自发反应信号,为VLA操作提供实时失败检测与恢复,实现高精度干预和低延迟响应。
AI中文摘要:
视觉-语言-动作(VLA)策略能够实现多样的机器人操作,但在执行过程中可能失败且无法识别自身错误。人类观察者提供了互补信号,因为意外的机器人行为在失败完成之前可能触发快速的语音、面部或言语反应。我们引入了SocialVLA,一个本地的、策略无关的社会感知网关,将自发性的人类反应转化为VLA操作的运行时干预信号。SocialVLA结合了因果副语言音频检测、视觉反应识别、明确的停止短语和机器人相关性估计。一种异步首事件融合机制从最早足够自信的信号触发VLA暂停,而一个独立的语音通道捕获言语纠正,用于参与者导向的继续、重启或指令修订。我们在物理Unitree G1操作上使用15名参与者评估了SocialVLA,包含238个注释的干预值得事件和1.038小时的非干预行为。冻结的离线重放实现了54.6%的召回率和69.5%的精确率,而未过滤的音频-视频融合达到64.3%的召回率。相关性估计将误停事件从100减少到57,并将精确率从60.5%提高到69.8%。在对未见过的第16名参与者的前瞻性部署中,冻结系统实现了59.5%的召回率和91.7%的精确率。检测器到融合的中位延迟为47.9毫秒,VLA门到物理暂停的延迟为336毫秒,反应起始到暂停的延迟为1.021秒。这些结果展示了从自发性社会反应到物理VLA中断和参与者导向恢复的完整本地路径。
英文摘要:
Vision-language-action (VLA) policies enable diverse robotic manipulation but can fail during execution without recognizing their own errors. Human observers provide complementary signals, as unexpected robot behavior can trigger rapid vocal, facial, or verbal reactions before failure is completed. We introduce SocialVLA, a local, policy-agnostic social perception gateway that converts spontaneous human reactions into runtime intervention signals for VLA manipulation. SocialVLA combines causal paralinguistic audio detection, visual reaction recognition, explicit stop phrases, and robot-relevance estimation. An asynchronous first-event fusion mechanism triggers a VLA hold from the earliest sufficiently confident signal, while a separate speech channel captures verbal corrections for participant-directed continuation, restart, or instruction revision. We evaluate SocialVLA on physical Unitree G1 manipulation using 15 participants, with 238 annotated intervention-worthy episodes and 1.038 h of non-intervention behavior. Frozen offline replay achieves 54.6% recall and 69.5% precision, while unfiltered audio-video fusion reaches 64.3% recall. Relevance estimation reduces false-stop episodes from 100 to 57 and increases precision from 60.5% to 69.8%. In prospective deployment on an unseen 16th participant, the frozen system achieves 59.5% recall and 91.7% precision. Median detector-to-fusion latency is 47.9 ms, VLA-gate-to-physical-hold latency is 336 ms, and reaction-onset-to-hold latency is 1.021 s. These results demonstrate a complete local pathway from spontaneous social reaction to physical VLA interruption and participant-directed recovery.