RobotEQ-Video:基于世界状态分类法的以视频为中心的社交主动智能基准
RobotEQ-Video: A Video-Centric Benchmark for Social Proactive Intelligence with World-State Taxonomy
- State Key Laboratory of Autonomous Intelligent Unmanned Systems, Tongji University(同济大学自主智能无人系统全国重点实验室)
- Tsinghua University(清华大学)
- The Chinese University of Hong Kong(香港中文大学)
- Shenzhen MSU-BIT University(深圳北理莫斯科大学)
- University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出RobotEQ-Video,一个以视频为中心的社交主动智能基准,通过四级世界状态分类法覆盖6领域、20维度、142+816个属性,含2000余视频和10万余标注,评估发现现有系统未达人类水平,推动SPI从静态图像迈向动态视频研究。
AI中文摘要:
社交主动智能(SPI)将主动辅助的范畴从任务完成度扩展到考虑多样化具身场景中的社交适当性。然而,先前的SPI研究面临两个关键局限。首先,现有工作聚焦于静态图像,而动态视频为推断人类状态和需求提供了关键线索,比孤立图像蕴含更丰富的信息。其次,先前工作常依赖自由形式的数据收集流程,无法保证对多样化场景的全面覆盖。为解决这些不足,我们提出RobotEQ-Video,将焦点从以图像为中心转向以视频为中心的分析。为确保视频覆盖的全面性,我们构建了一个分层世界状态分类法,组织为四级由粗到细的结构,包含6个领域、20个维度、142个一级属性和816个二级属性。由此产生的基准包含2000多个视频、10万余条人工标注以及16000多个用于评估行为恰当性的标签。基准评估显示,当前系统仍不可靠,且未达到人类表现水平。我们进一步探索了世界模型如何帮助解决这一任务。这项工作将SPI研究从静态图像推进到动态视频,并确保基准测试中场景覆盖更加全面。
英文摘要:
Social Proactive Intelligence (SPI) extends proactive assistance beyond task completeness to consider social appropriateness in diverse embodied scenarios. However, prior SPI research faces two key limitations. First, existing work focuses on static images, whereas dynamic videos provide crucial cues for inferring human states and needs, offering richer information than isolated images. Second, prior work often relies on free-form data collection pipelines, which fail to guarantee comprehensive coverage of diverse scenarios. To address these gaps, we introduce RobotEQ-Video, shifting the focus from image-centric to video-centric analysis. To ensure comprehensive video coverage, we construct a hierarchical world-state taxonomy organized into a four-level coarse-to-fine structure, comprising 6 domains, 20 dimensions, 142 level-1 attributes, and 816 level-2 attributes. The resulting benchmark comprises 2K+ videos with 100K+ human annotations and 16K+ labels for assessing behavior properness. Benchmark evaluation reveals that current systems remain unreliable and fall short of human performance. We further explore how world models can help tackle this task. This work advances SPI research from static images to dynamic videos and ensures more comprehensive scenario coverage during benchmarking.