arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FineHOI:面向零样本人-物交互检测的部件感知密集表示

FineHOI: Part-Aware Dense Representations for Zero-Shot Human-Object Interaction Detection

Francesco Tonini, Lorenzo Vaquero, Mohammad Mahdi Derakhshani, Cees Snoek, Elisa Ricci, Cigdem Beyan

arXiv 2609.05959首次发表:更新:

发表机构

University of Trento; Fondazione Bruno Kessler; University of Amsterdam; University of Verona(特伦托大学; 布鲁诺·凯斯勒基金会; 阿姆斯特丹大学; 维罗纳大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FineHOI提出部件感知密集表示框架,通过自适应部件级注意力与区域感知交互Transformer,提升零样本人-物交互检测性能,尤其在未见交互上表现突出。

AI 中文摘要

人-物交互(HOI)检测旨在定位图像中的人与物体,并对其交互进行分类。零样本HOI聚焦于识别训练过程中未观察到的交互,要求模型能够泛化到未见过的动词-物体组合。近期方法利用视觉-语言模型(VLMs),受益于其丰富的语义表示。然而,这些方法往往依赖全局或检测器中心特征,这会压缩交互线索并阻碍细粒度的空间推理。为克服这一局限,我们提出FineHOI,一种从密集块级特征显式建模交互的零样本HOI框架。我们的方法受以下观察启发:人-物交互由局部空间关系定义,而全局和检测器中心表示无法保留这些关系。为此,我们引入自适应部件级注意力模块,通过无监督聚类将人与物体分解为语义连贯的部件,并根据其交互相关性重新加权。这些表示随后通过区域感知交互Transformer进行整合,该Transformer融合部件感知特征与全局特征,生成最终的HOI嵌入。大量实验表明,FineHOI持续优于现有零样本HOI方法,尤其在未见交互上取得显著提升。代码可在该https URL获取。

英文摘要

Human-Object Interaction (HOI) detection aims to localize humans and objects in images and classify their interactions. Zero-shot HOI focuses on recognizing interactions that are not observed during training, requiring models to generalize beyond seen verb-object compositions. Recent approaches leverage Vision-Language Models (VLMs), benefiting from rich semantic representations. However, they often rely on global or detector-centric features that compress interaction cues and hinder fine-grained spatial reasoning. To overcome this limitation, we propose FineHOI, a zero-shot HOI framework that explicitly models interactions from dense patch-level features. Our approach is motivated by the observation that human-object interactions are defined by localized spatial relationships, which are not preserved by global and detector-centric representations. To this end, we introduce an Adaptive Part-Level Attention module that decomposes humans and objects into semantically coherent parts via unsupervised clustering, and re-weights them based on their interaction relevance. These representations are then integrated through a Region-Aware Interaction Transformer that integrates part-aware and global features and produces the final HOI embedding. Extensive experiments demonstrate that FineHOI consistently outperforms existing zero-shot HOI methods, achieving particularly strong gains on unseen interactions. Code is available at https://github.com/francescotonini/fine-hoi.

CommentsAccepted at ACM Multimedia 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑