身份感知的人-物交互动作描述
Identity-Aware Human-Object Interaction Motion Captioning
AI总结:
针对现有HOI动作描述未关联主体身份的局限,提出身份感知HOI动作描述任务,构建ID-HOINet模型,含MVIML与TSCR组件,在BEHAVE和InterCap数据集上实现SOTA性能。
AI中文摘要:
现有的人-物交互(HOI)动作描述方法通常使用“一个人”或“某人”等通用术语描述发生的事件,未将描述与主体身份关联。为解决这一局限,我们提出身份感知的人-物交互动作描述任务,该任务要求每个生成的描述同时指定主体身份和对应的HOI动作,例如模型生成“Sub_ID 抬起椅子”而非“一个人抬起椅子”。针对该任务,我们基于 BEHAVE 和 InterCap 数据集设计身份感知的 HOI 动作描述方法,进一步提出 ID-HOINet,该模型从多视角视频中学习,同时支持单视角身份感知的 HOI 动作描述生成。ID-HOINet 包含两个核心组件:多视角身份-动作学习模块(MVIML)和两阶段描述重写策略(TSCR)。MVIML 通过建模时间阶段和相机视角间的依赖关系,从多视角视频中学习,捕获身份和交互动作特征。在推理阶段,TSCR 首先检索主体身份并生成与身份无关的 HOI 动作描述,随后用预测的身份重写这些描述,生成最终的身份感知 HOI 动作描述。实验表明,ID-HOINet 达到了最先进的性能,代码将在论文接收后发布。
英文摘要:
Existing human-object interaction (HOI) motion captioning methods typically describe what happens while referring to the subject using generic terms such as "a person" or "someone", without grounding the caption in subject identity. To address this limitation, we introduce Identity-Aware Human-Object Interaction Motion Captioning task. This task requires each generated caption to specify both the subject identity and the corresponding HOI motion. For example, the model generates "Sub_ID lifts the chair" rather than "A person lifts the chair". For this task, we design identity-aware HOI motion captions based on the BEHAVE and InterCap datasets. We further propose ID-HOINet, which learns from multi-view videos while supporting single-view identity-aware HOI motion caption generation. ID-HOINet contains two core components: Multi-View Identity-Motion Learning Module (MVIML) and Two-Stage Caption Rewriting Strategy (TSCR). MVIML learns from multi-view videos by modeling dependencies across temporal stages and camera viewpoints, capturing identity and interaction motion features. At inference, the TSCR first retrieves the subject identity and generates identity-agnostic HOI motion captions. TSCR then rewrites these captions with the predicted identity to produce the final identity-aware HOI motion captions. Experiments demonstrate that ID-HOINet achieves state-of-the-art performance. Code will be released upon acceptance.