arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

开放词汇注视对象预测:基准与方法

Open-Vocabulary Gaze Object Prediction: Benchmark and Method

Binglu Wang, Sensen Niu, Ying Chen, Guangyu Guo

arXiv 2607.18827首次发表:更新:

发表机构

Xi’an University of Architecture and Technology; DAMO Academy, Alibaba Group; Zhejiang University(西安建筑科技大学; 阿里巴巴达摩院; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究注视对象预测问题,提出基于文本驱动对象发现和注视引导选择模块的框架,并引入梯度信息选择调整,构建含86个自然类别的基准,所提模型在开放和封闭词汇设置中均表现出色。

AI 中文摘要

注视对象预测(GOP)旨在定位和识别人类注视的对象,这对理解以人为中心的交互至关重要。现有方法通常在封闭词汇范式下训练,标签空间固定且在特定场景数据集上评估,限制了其在现实场景中的适用性。为此,我们引入了用于注视对象预测的多样场景(DiSG)基准,包含86个自然类别,便于评估开放词汇GOP(OVGOP)。在此基础上,我们提出一个框架,利用文本驱动的对象发现来定位潜在注视候选对象,并通过注视引导选择模块从候选对象中确定目标对象。此外,为更好地捕捉不同自然类别间的语义知识,我们引入梯度信息选择调整(GIST)来选择性更新与给定类别词汇最相关的参数。大量实验表明,我们提出的模型在开放词汇设置中有效,且在传统封闭词汇设置中也优于现有方法。基准和代码可通过此https链接获取。

英文摘要

Gaze Object Prediction (GOP) aims to localize and recognize the objects humans attend to, a task crucial for understanding human-centric interactions. However, existing methods are typically trained under a closed-vocabulary paradigm with a fixed label space and evaluated on scene-specific datasets, limiting their applicability to real-world scenarios where gaze targets often follow a long-tail distribution or belong to unseen categories. To address this gap, we introduce Diverse Scenes for Gaze object prediction (DiSG), a benchmark containing 86 in-the-wild categories that facilitates the evaluation of Open-Vocabulary GOP (OVGOP). Building on DiSG, we propose a framework that leverages text-driven object discovery to localize potential gaze candidates, with a gaze-guided selection module to pinpoint the intended target from the candidate objects. Furthermore, to better capture semantic knowledge across diverse in-the-wild categories, we introduce Gradient-Informed Selection Tuning (GIST) to selectively update parameters most relevant to a given class vocabulary. Extensive experiments demonstrate that our proposed model performs effectively in open-vocabulary settings and also outperforms existing methods in the conventional closed-vocabulary setting. The benchmark and code is available at https://github.com/sensniu/ovgop.

CommentsAccepted at ACM Multimedia 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑