arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

社会直觉 vs. 机器推理:从多模态预测人机交互

Social Intuition vs. Machine Reasoning: Anticipating Human-Robot Interaction from multiple modalities

Raphael Lorenzo-Louis, Bertrand Luvison, Serena Ivaldi

arXiv 2609.07394首次发表:更新:

发表机构

Inria; CNRS; Loria; HUCEBOT; CEA, List; Université Paris-Saclay(法国国家信息与自动化研究所; 法国国家科学研究中心; 洛里亚实验室; HUCEBOT; 法国原子能和替代能源委员会 技术与信息技术部门; 巴黎-萨克雷大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过对比人类与轻量级姿态模型及视觉-语言模型在HUI360数据集上的交互预测表现,发现人类在完整视频输入下显著优于机器,表明推理能力不足以匹敌社会直觉。

AI 中文摘要

从自身视角预测一个人是否会进行交互,对人类来说是一项高度直觉性的任务,它依赖于多种线索的组合。我们研究了人类在仅使用姿态或完整视频输入的情况下,从服务机器人的视角预测人的交互意图的表现,然后对不同的轻量级基于姿态的模型和最先进的视觉-语言模型进行了基准测试。我们在HUI360数据集上进行了基准测试,使用了一个固定的试点子集,包含100条测试轨迹(25条正样本,75条负样本)。我们发现,在仅使用姿态输入时,人类标注者的表现优于经过训练的轻量级姿态模型,但优势不大(F1分数提高0.08)。然而,当提供带有目标边界框的完整自我中心视频时,人类标注者的表现显著更好,并大幅超越视觉-语言模型(F1分数提高0.2)。我们还比较了不同大小和不同输入条件下的视觉-语言模型,发现最佳结果与模型大小无关。我们的结果证实,预测交互对社交机器人来说是一项具有挑战性的任务,具备推理能力的模型是必要的,但仅凭其实际推理能力不足以匹敌人类的社会直觉。

英文摘要

Anticipating whether a person will interact from one's own perspective is a highly intuitive task for humans, that relies on a combination of cues. We investigate how humans perform at predicting a person's intention to interact from a service robot's point of view, using pose-only or full video input, then benchmark different lightweight pose-based models and state-of-the-art vision-language models. We conducted our benchmark on the HUI360 dataset on a fixed pilot subset of 100 test tracks (25 positive, 75 negative). We found that with pose-only input, human annotators outperform lightweight trained pose models but not by large margins (+0.08 in F1-Score). But when given full egocentric video with a target bounding box, human annotators perform substantially better and largely outperform the Vision-Language Models (+0.2 in F1-Score). We also compared VLMs of different size and under different input conditions, and found that the best results do not correlate with model size. Our result confirms that predicting interactions is a challenging task for social robots and that reasoning-capable models are necessary but their actual reasoning capabilities alone do not suffice to match the social intuition of humans.

Comments5 pages, 4 figures. Workshop paper accepted to The 4th Workshop on Nonverbal Cues for Human-Robot Cooperative Intelligence at IROS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑