发表机构
Shanghai Jiao Tong University; Xinjiang University; Beijing Jiaotong University; Shanghai AI Laboratory(上海交通大学; 新疆大学; 北京交通大学; 上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对免训练人-物交互检测中推理碎片化和语义循环问题,提出HarnessHOI框架,通过交互引导感知和几何裁决模块,在HICO-DET和V-COCO上达到免训练方法最优性能。
AI 中文摘要
人-物交互(HOI)检测旨在定位人-物对并识别其交互行为。传统的监督方法性能强大,但依赖于特定任务的训练。近年来的多模态大语言模型(MLLMs)凭借其广泛的视觉-语义知识和多功能的感知与推理能力,为免训练的HOI检测提供了一条有前景的途径。然而,现有方法大多通过松散协调的推理阶段来调用这些能力。这种碎片化的执行限制了交互假设在引导视觉探索中的作用,导致关键参与者被忽略,局部歧义无法解决。此外,将早期的语义假设传播到后续的视觉定位和关系预测中,会引发自我强化的语义循环。为解决这些挑战,我们提出了HarnessHOI,一个免训练框架,将被动的MLLM推理转变为主动的以交互为中心的利用机制。具体而言,我们引入了一种交互引导的感知机制,该机制将新兴的交互假设投影回视觉空间,以发现缺失的参与者,并通过有针对性的观察来细化模糊的证据。此外,一个与关系无关的几何裁决模块协调多源证据,为跨多个动作和语义角色的定位交互推理建立统一的视觉空间基础。在HICO-DET和V-COCO上的大量实验表明,HarnessHOI在免训练方法中达到了最先进的性能,证实了所提出的利用机制在复杂交互理解中的有效性。代码将在发表后发布。
英文摘要
Human-object interaction (HOI) detection aims to localize human-object pairs and recognize their interactions. Traditional supervised methods perform strongly but rely on task-specific training. Recent multimodal large language models (MLLMs) offer a promising route to training-free HOI detection through their broad visual-semantic knowledge and versatile perceptual and reasoning capabilities. However, existing approaches largely invoke these capabilities through loosely coordinated inference stages. This fragmented execution restricts the role of interaction hypotheses in guiding visual exploration, leaving key participants overlooked and local ambiguities unresolved. Furthermore, propagating early semantic assumptions through subsequent visual grounding and relation prediction induces self-reinforcing semantic circularity. To resolve these challenges, we propose HarnessHOI, a training-free framework that transforms passive MLLM inference into an active interaction-centric harness. Specifically, we introduce an interaction-guided perception mechanism that projects emerging interaction hypotheses back into the visual space to discover missing participants and refine ambiguous evidence through targeted observation. Furthermore, a relation-agnostic geometric adjudication module reconciles multi-source evidence to establish a unified spatial basis for grounded interaction reasoning across multiple actions and semantic roles. Extensive experiments on HICO-DET and V-COCO demonstrate that HarnessHOI achieves state-of-the-art performance among training-free methods, confirming the effectiveness of the proposed harness for complex interaction understanding. Code will be released upon publication.