释放多模态大语言模型用于野外无训练的人机交互检测
Unleashing Multimodal Large Language Models for Training-free HOI Detection in the Wild
- Wangxuan Institute of Computer Technology, Peking University(北京大学王选计算机技术研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对传统人机交互检测方法在开放世界和组合场景中泛化能力受限的问题,提出无训练的AgentHOI框架,利用基础模型多模态推理能力,通过上下文感知多轮推理和多方面交互定位机制,在实际场景中取得优于现有监督和弱监督方法的性能。
AI中文摘要:
传统上,人机交互检测(HOID)被视为针对预定义交互类别的监督检测问题。虽然这种范式在封闭集基准测试中表现出色,但它将交互理解与特定数据集的监督紧密相连,限制了其在开放世界和组合场景中的泛化能力。近期的HOI检测器试图通过提示策略利用多模态大语言模型(MLLMs)来转移特定交互知识。然而,这种基于提示的方法主要侧重于从预训练模型中提取判别性表示,而未充分探索其固有的多模态推理能力。因此,它们难以在模糊和开放世界的交互场景中提供信息丰富的上下文推理。在这项工作中,我们提出了AgentHOI,这是一个无训练的、具有智能体特性的框架,它将基础模型的通用多模态推理能力转移到野外的HOI检测中。AgentHOI不是学习交互分类器,而是模块化地协调互补的视觉基础模块,以协同方式执行开放式语义推理和空间定位。为了应对复杂场景中交互发现不完整和定位模糊的挑战,我们引入了两个关键机制:(1)上下文感知多轮推理,逐步完善交互假设,以确保全面和组合式的HOI发现;(2)多方面交互定位,通过生成整合语义、空间和外观线索的特定实例描述来提高定位精度。大量实验表明,AgentHOI在实际场景中比最先进的监督和弱监督方法具有更高的性能,尽管它无需HOID数据进行训练。
英文摘要:
Human-object interaction detection (HOID) has traditionally been formulated as a supervised detection problem over predefined interaction categories. While such paradigms achieve strong performance on closed-set benchmarks, they fundamentally entangle interaction understanding with dataset-specific supervision, limiting their ability to generalize to open-world and compositional scenarios. Recent HOI detectors attempt to leverage MLLMs through prompting strategies to transfer interaction-specific knowledge. However, such prompt-based approaches primarily focus on extracting discriminative representations from pretrained models, while underexploring their inherent multimodal reasoning capabilities. As a result, they struggle to provide informative contextual reasoning for ambiguous and open-world interaction scenarios. In this work, we present AgentHOI, a training-free, agentic framework that transfers the generalist multimodal reasoning capabilities of foundation models to HOI detection in the wild. Instead of learning interaction classifiers, AgentHOI modularly orchestrates complementary vision foundation modules to perform open-ended semantic reasoning and spatial grounding in a coordinated manner. To address the challenges of incomplete interaction discovery and ambiguous localization in complex scenes, we introduce two key mechanisms: (1) Context-aware Multi-round Reasoning, which progressively refines interaction hypotheses to ensure exhaustive and compositional HOI discovery, and (2) Multifaceted Interaction Localization, which enhances grounding precision by generating instance-specific descriptions that integrate semantic, spatial, and appearance cues. Extensive experiments demonstrate that AgentHOI achieves superior performance over state-of-the-art supervised and weakly supervised methods in real-world settings, despite requiring no HOID data for training.