发表机构
Karlsruhe Institute of Technology; TED University(卡尔斯鲁厄理工学院; 泰德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究任务驱动的视觉显著度预测问题,提出基于自然语言任务描述调节视觉注意力的模型,生成依赖任务的显著度图,经分析证明纳入任务语义可更忠实地建模目标导向视觉注意力。
AI 中文摘要
视觉显著度旨在预测图像中最能吸引人类视觉注意力的区域。大多数显著度模型假设自由观看条件,而人类注意力常受明确任务目标影响。本文提出一个基于自然语言任务描述来调节视觉注意力的模型,以解决任务驱动的显著度预测问题。该模型生成依赖任务的显著度图,反映不同观看意图下注意力的转移。通过定量和定性分析表明,纳入明确任务语义能更忠实地对目标导向的视觉注意力进行建模。
英文摘要
Visual saliency aims to predict the regions of an image most likely to attract human visual attention. While most saliency models assume free-viewing conditions, human attention is often shaped by explicit task goals. In this work, we address task-driven saliency prediction by proposing a model that conditions visual attention on natural-language task descriptions. The model produces task-dependent saliency maps that reflect how attention shifts under different viewing intents. Through quantitative and qualitative analysis, we show that incorporating explicit task semantics enables more faithful modeling of goal-directed visual attention.
Comments15 pages, 5 figures