arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31364cs.CV

OpenVAM:基于视觉语言模型的开世界视觉注意力建模

OpenVAM: Open-World Visual Attention Modeling with VLMs

  • University of Tehran(德黑兰大学)
  • University of Toronto(多伦多大学)
  • Vector Institute(向量研究所)
  • University Health Network(大学健康网络)

机构由 AI 辅助整理,请以论文原文为准。

Kiana Hooshanfar, Amirhossein Kazerouni, Alireza Hosseini, Michael Brudno, Babak Taati

AI总结:

OpenVAM提出统一框架,结合密集视觉通路与视觉-语言语义头,实现跨领域注视预测的通用性与可解释性,并通过三阶段训练和可扩展注释流水线提升鲁棒性。

AI中文摘要:

预测人类注视是众多应用的核心能力,这些应用涵盖从网页/UI设计分析到机器人和人机交互等领域。然而,大多数视觉注意力建模方法仅输出一个密集的显著性图,这通常不足以支撑行动:实践者需要将注意力峰值与场景中的离散元素(什么)联系起来,并理解这些峰值在上下文中的驱动因素(为什么),同时还要在自然图像、商业内容和UI/网页布局之间保持对领域偏移的鲁棒性。因此,我们提出了OpenVAM(基于视觉语言模型的开世界视觉注意力建模),这是一个统一框架,共同解决跨异构领域(自然场景、商业图像和UI/网页布局)和监督模态的通用性与可解释性问题。OpenVAM采用了一种解耦但对齐的设计:一个专用的密集视觉通路提供稳定、空间精确的定位,而一个遵循指令的视觉-语言语义头则基于相同图像和数据类型提示生成有根据的“什么/为什么”解释。三阶段训练策略保留了强大的定位先验,同时通过参数高效适配逐步引入语言接地并改善解释对齐,而不干扰显著性分支。我们进一步提出了一种可扩展的流水线,用于生成多领域显著性推理注释,以用于训练和系统评估。跨多个数据集的实验表明,OpenVAM在领域偏移下提高了鲁棒性,同时生成基于图像的解释,使显著性预测更具可解释性。

英文摘要:

Predicting human gaze is a core capability for applications ranging from web/UI design analysis to robotics and human-computer interaction. Yet, most visual attention modeling methods output only a dense saliency map, which is often insufficient for action: practitioners need to connect attention peaks to discrete elements in the scene (what) and understand the drivers of those peaks in context (why), while remaining robust to domain shift across natural images, commercial content, and UI/web layouts. We, therefore, introduce OpenVAM (Open-world Visual Attention Modeling with VLMs), a unified framework that jointly addresses universality and explainability across heterogeneous domains (natural scenes, commercial imagery, and UI/web layouts) and supervision modalities. OpenVAM adopts a decoupled-but-aligned design: a dedicated dense visual pathway provides stable, spatially precise localization, while an instruction-following vision--language semantic head generates grounded what/why explanations conditioned on the same image and data-type prompts. A three-stage training strategy preserves strong localization priors while progressively introducing language grounding and improving explanation alignment via parameter-efficient adaptation without perturbing the saliency branch. We further propose a scalable pipeline to generate multi-domain saliency-reason annotations for training and systematic evaluation. Experiments across diverse datasets show that OpenVAM improves robustness under domain shift while producing image-grounded explanations that make saliency predictions more interpretable.

↑