发表机构
University of Illinois Urbana-Champaign; Google(伊利诺伊大学厄巴纳-香槟分校; 谷歌公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有注视估计方法依赖中间阶段、缺乏灵活性的局限,提出PGE任务,构建Gaze-Co数据集,开发GazeAnywhere模型,实现端到端的概念驱动注视目标估计并达到最优性能。
AI 中文摘要
从野外图像中估计人类注视目标是一项重要且极具挑战性的任务。现有方法主要采用脆弱的多阶段流程,需要明确的输入(如头部边界框和人体姿态)来识别注视分析的主体,因此检测误差会级联并导致任务失败。此外,这些现有工作缺乏通过自然语言提示指定注视分析任务的灵活性,而该方法已被证明在其他图像分析任务中具有显著的便利性和可扩展性优势。为克服这些局限,我们提出了可提示注视目标估计(Promptable Gaze Target Estimation, PGE)任务,这是一种新的、端到端的、以概念为驱动的注视分析范式。PGE将注视预测建立在灵活的用户文本或视觉提示(例如“穿红色衬衫的男孩”或“点[0.52, 0.48]处的人”)之上,以识别注视分析的特定主体。该方法将主体定位与注视估计相结合,消除了对中间分析阶段的刚性依赖。我们开发了一个可扩展的数据引擎,生成了Gaze-Co(Gaze Estimation with Concepts),这是一个包含12万对高质量、带提示标注的图像的数据集和基准。我们还提出了GazeAnywhere,这是首个专为PGE设计的模型。GazeAnywhere使用基于Transformer的检测器融合来自冻结编码器的特征,同时解决主体定位、是否在帧内存在以及注视目标热图估计问题。GazeAnywhere在多个PGE基准上实现了最先进的性能,即使在困难的域外真实世界临床数据集上也为这一新问题设定了强大的基线。GazeAnywhere已开源,网址为this http URL。
英文摘要
Estimating human gaze targets from images in-the-wild is an important and formidable task. Existing approaches primarily employ brittle, multi-stage pipelines that require explicit inputs, like head bounding boxes and human pose, in order to identify the subject of gaze analysis. As a result, detection errors can cascade and lead to failure. Moreover, these prior works lack the flexibility of specifying the gaze analysis task via natural language prompting, an approach which has been shown to have significant benefits in convenience and scalability for other image analysis tasks. To overcome these limitations, we introduce the Promptable Gaze Target Estimation (PGE) task, a new end-to-end, concept-driven paradigm for gaze analysis. PGE conditions gaze prediction on flexible user text or visual prompts (e.g., "the boy in the red shirt" or "person in point [0.52, 0.48]") to identify a specific subject for gaze analysis. This approach integrates subject localization with gaze estimation, and eliminates the rigid dependency on intermediate analysis stages. We develop a scalable data engine to generate Gaze-Co (Gaze Estimation with Concepts), a dataset and benchmark of 120K high-quality, prompt-annotated image pairs. We also propose GazeAnywhere, the first model designed for PGE. GazeAnywhere uses a transformer-based detector to fuse features from frozen encoders and simultaneously solves subject localization, in/out-of-frame presence, and gaze target heatmap estimation. GazeAnywhere achieves state-of-the-art performance on multiple PGE benchmarks, setting a strong baseline for this new problem even on a difficult out-of-domain, real-world clinical dataset. GazeAnywhere is open-sourced in github.com/IrohXu/GazeAnywhere.
CommentsCVPR 2026 Code and Benchmark are aviliable at https://github.com/IrohXu/GazeAnywhere and https://huggingface.co/datasets/IrohXu/Gaze-Co-Benchmark