发表机构
Université de Montréal; Mila – Quebec AI Institute; McGill University; McMaster University; Meta Platforms(蒙特利尔大学; 米拉-魁北克人工智能研究所; 麦吉尔大学; 麦克马斯特大学; Meta平台公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ReactHuman是首个物理基础的类人反应决策基准,通过17个事件族和1000多个可复现场景评估多模态大语言模型在突发家庭危险中的即时反应能力,发现现有模型在反应安全性上远未解决,且失败不随规模缩小。
AI 中文摘要
对突发物理危险的响应(接住滑落的盘子、躲避落下的刀具)既是对具身智能的有意义测试,也是将多模态大语言模型(MLLMs)作为家用机器人决策核心的硬性要求。然而,现有评估要么通过视频问答被动地探测直觉物理,要么针对导航和重排等深思熟虑的长期任务;没有一项评估能衡量模型能否将物理理解转化为即时的、安全关键的行动。我们提出了ReactHuman,这是首个用于类人反应决策的物理基础基准,其中被评估的MLLM充当面对突发家庭危险的模拟人形机器人的大脑;它涵盖17个事件族和超过1000个逐位可复现的场景,具有来自240Hz刚体模拟的精确、无注释的真实数据,包括外观与其物理特性相矛盾的对抗性物体(泡沫铁砧、钢制苹果)。我们进一步设计了一个五指标套件,从三个维度对每个反应进行评分:合理性、安全性和物理基础性。我们物理执行每个承诺的计划,使决策具有可观察的后果。利用这一框架,我们评估了七个代表性的MLLM。结果表明,反应安全性远未解决:模型大约每三个危险中就有一个处理不当,根据固定倾向而非观察到的场景行动,信任外观而非运动,即使所选动作正确,也会在米级尺度上错过拦截点;这些失败都不会随模型规模而缩小。因此,ReactHuman为物理基础、安全意识的具身智能体提供了细粒度的诊断和可扩展的训练信号。基准可在此处找到:此https URL
英文摘要
Reacting to sudden physical hazards (catching a slipping plate, dodging a falling knife) is both a meaningful test of embodied intelligence and a hard requirement for deploying multimodal large language models (MLLMs) as the decision coreof household robots. Existing evaluations, however, probe intuitive physics passively through question answering over videos, or target deliberate, long-horizon tasks such as navigation and rearrangement; none measure whether a model can turn physical understanding into immediate, safety-critical action. We introduce ReactHuman, the first physics-grounded benchmark for human-like reactive decision-making, in which the evaluated MLLM acts as the brain of a simulated humanoid facing sudden household hazards; it spans 17 event families and over 1,000 bit-for-bit reproducible scenes with exact, annotation-free ground truth derived from 240 Hz rigid-body simulation, including adversarial objects whose appearance contradicts their physics (a foam anvil, a steel apple). We further design a five-metric suite that scores each reaction along three axes: reasonable, safe, and physically grounded. We physically execute every committed plan so that decisions have observable consequences. With this harness we evaluate seven representative MLLMs. Results show that reactive safety is far from solved: models mishandle roughly one hazard in three, act from fixed dispositions rather than the observed scene, trust appearance over motion, and miss interception points at meter scale even when the chosen action is correct; none of these failures shrink with model scale. ReactHuman thus offers both a fine-grained diagnosis and a scalable training signal toward physically grounded, safety-aware embodied agents. The benchmark can be found here: https://huggingface.co/datasets/Alan123/reacthuman-benchmark-scaled