发表机构
Karlsruhe Institute of Technology (KIT); Institute for Computer Science, Artificial Intelligence and Technology (INSAIT); Hunan University; University of Bremen; ETH Zurich(卡尔斯鲁厄理工学院; 计算机科学、人工智能与技术研究所; 湖南大学; 不来梅大学; 苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出HEIR基准和CoRISP方法,用于学习人-实体交互中的完整参与者-角色集,通过共享实体身份和角色条件证据预测,在多个基准上显著提升事件组合评估性能。
AI 中文摘要
理解人-实体交互需要恢复每个人-动作事件的参与者、角色和共享身份。这种结构可以通过明确谁作用于哪些实体以及如何作用,来支持具身智能体,为共享环境中的预期和协调提供信息。标准的HOI指标对单个链接进行评分,导致完整的事件组合未被充分衡量。我们引入了HEIR(具有功能角色的人-实体交互),一个用于跨物体、人际和自指向交互的完整接地参与者-角色集的图像基准。它包含18,730张图像、六个角色、105个动作和437个名词,具有共享实体、角色变化和重复填充;51.6%的图像包含多个参与者,62.1%包含多个动作。HEIR将关系AP与完整集AP和结构评估配对。我们还引入了CoRISP(组合式角色感知交互集预测),它使用共享实体身份来组合角色条件证据并预测归一化的参与者-角色集。基数和角色多重性势通过事件大小和角色组成耦合分配,并进行精确的逐事件归一化。在16个基线中,即使在调整动作权重后,关系和完整事件排名也会出现分歧。CoRISP在HEIR的重复角色事件和共享参与者图像上分别领先2.87和3.82个Set mAP点。在V-COCO上,CoRISP在场景1/2下的双槽动作中实现了73.72/76.23的角色AP和61.06/68.59的完整集AP。这些结果表明,在个体关系之外学习和评估事件组合的价值。代码和数据集在此https URL公开可用。
英文摘要
Understanding human-entity interactions requires recovering each person-action event's participants, roles, and shared identities. This structure can support embodied agents by clarifying who acts on which entities and how, informing anticipation and coordination in shared environments. Standard HOI metrics score individual links, leaving complete event composition undermeasured. We introduce HEIR (Human-Entity Interactions with Functional Roles), an image benchmark for complete grounded participant-role sets across object, interpersonal, and self-directed interactions. It contains 18,730 images, six roles, 105 actions, and 437 nouns, with shared entities, role changes, and repeated fillers; 51.6% of images contain multiple actors and 62.1% contain multiple actions. HEIR pairs relation AP with complete-set AP and structural evaluation. We also introduce CoRISP (Compositional Role-aware Interaction Set Prediction), which uses shared entity identities to combine role-conditioned evidence and predict normalized participant-role sets. Cardinality and role-multiplicity potentials couple assignments through event size and role composition, with exact per-event normalization. Across 16 baselines, relation and complete-event rankings diverge even after aligning action weights. CoRISP leads the evaluated systems on repeated-role events and shared-participant images in HEIR by 2.87 and 3.82 Set mAP points, respectively. On V-COCO, CoRISP achieves 73.72/76.23 role AP and 61.06/68.59 complete-set AP on two-slot actions under Scenarios 1/2. These results show the value of learning and evaluating event composition alongside individual relations. The code and dataset are publicly available at https://github.com/Kratos-Wen/HEIR.
Comments24 pages, 4 figures. Code and dataset: https://github.com/Kratos-Wen/HEIR