HOIBlender:融合轻量级检测与视觉-语言先验的高效人-物交互检测
HOIBlender: Blending Lightweight Detection with Vision-Language Priors for Efficient Human-Object Interaction Detection
浏览论文内容
中文总结 AI 辅助
HOIBlender提出一种高效人-物交互检测方法,通过轻量级解码流程融合检测器接地视觉标记、空间推理和BLIP-2语义先验,在HICO-DET上以9个训练周期达到44.49 mAP,超越现有方法。
中文摘要 AI 辅助
人-物交互(HOI)检测需要定位交互的人-物对并识别连接它们的动词,这通常面临严重的长尾监督问题。近期方法通过更强的检测器和视觉-语言先验提高了准确性,但许多方法仍在检测器之上堆叠重型Transformer编码器、复杂的去噪调度或事后语义校准。我们提出\ extbf{HOIBlender},一种高效的HOI检测器,其命名源于其核心设计原则:在单一轻量级解码流程中融合检测器接地视觉标记、空间主体-客体推理和BLIP-2语义先验。HOIBlender基于RF-DETR/LW-DETR风格的基础架构,采用DINOv2骨干网络,直接从多尺度投影器中选择Top-$K$图像条件标记作为主体和客体候选,去除了先前HOI方法保留的专用编码器阶段。双阶段解码器首先稳定人-物几何关系,然后通过渐进式BLIP-2先验融合执行动词和HOI分类,分类器权重从BLIP-2文本嵌入初始化以处理长尾类别。分组查询训练进一步丰富了优化过程而不增加推理成本。在三种模型规模(Nano、Small、2XL)下,HOIBlender在HICO-DET上持续优于SOV-STG-VLA和Hybrid-SOV-VLA,仅用$9$个训练周期即达到$44.49$的Default Full mAP,同时保持有竞争力的延迟和参数预算。这些结果表明,轻量级检测、结构化空间-语义解码和深度集成的视觉-语言先验可以融合到单一高效的HOI流程中。
英文摘要
Human-object interaction (HOI) detection requires grounding an interacting human-object pair and recognizing the verb that links them, often under severe long-tail supervision. Recent methods improve accuracy with stronger detectors and vision-language priors, but many still stack heavy transformer encoders, intricate denoising schedules, or post-hoc semantic calibration on top of the detector. We present \textbf{HOIBlender}, an efficient HOI detector named after its core design principle: blending detector-grounded visual tokens, spatial subject-object reasoning, and BLIP-2 semantic priors inside one lightweight decoding pipeline. HOIBlender builds on an RF-DETR/LW-DETR-style foundation with a DINOv2 backbone and selects top-$K$ image-conditioned tokens directly from the multi-scale projector as subject and object candidates, removing the dedicated encoder stage retained by prior HOI methods. A dual-stage decoder first stabilizes human-object geometry and then performs verb and HOI classification through progressive BLIP-2 prior fusion, with classifier weights initialized from BLIP-2 text embeddings for long-tail categories. Grouped-query training further enriches optimization without increasing inference cost. Across three model scales (Nano, Small, 2XL), HOIBlender consistently outperforms SOV-STG-VLA and Hybrid-SOV-VLA on HICO-DET, reaching $44.49$ Default Full mAP in only $9$ training epochs while maintaining competitive latency and parameter budgets. These results show that lightweight detection, structured spatial-semantic decoding, and deeply integrated vision-language priors can be blended into a single efficient HOI pipeline.
发表机构
- The University of Electro-Communications(电气通信大学)
机构由 AI 辅助整理,请以论文原文为准。