PRE-MAP:用于高分辨率多属性点预测的个性化强化眼动多模态大语言模型
PRE-MAP: Personalized Reinforced Eye-tracking Multimodal LLM for High-Resolution Multi-Attribute Point Prediction
- Jilin University(吉林大学)
- Peking University(北京大学)
- Mininglamp Technology(Mininglamp科技)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对个性化注意与高分辨率多点定位难题,论文构建SPA-ADV眼动数据集,并提出基于MLLM和C-GRPO强化优化的PRE-MAP模型。
AI中文摘要:
由个体偏好驱动的视觉选择性注意,通过连接主观认知机制与客观视觉元素,调节人类对视觉刺激的优先级排序,从而引导动态视觉场景的语义解释与层级化处理。然而,现有模型和数据集大多忽视主观认知多样性对注视行为的影响。传统显著性预测模型通常采用分割方法,依赖低分辨率图像生成显著性热图,随后再上采样至原生分辨率,这限制了其捕捉个性化注意模式的能力。此外,MLLM受到幻觉等因素制约,在涉及多点预测的任务中严格遵循预期格式的成本很高,且实现精确的点定位颇具挑战。为解决上述局限,我们提出了面向广告视频的主观个性化注意数据集,即SPA-ADV,这是一个大规模多模态数据集,通过486个视频记录了4500多名年龄和性别各异的参与者的注视行为。此外,我们提出PRE-MAP,这是一种新型眼动显著性模型,通过强化学习优化的眼动来刻画个性化视觉差异;该模型构建于MLLM之上,并在多属性用户画像引导下预测点。为确保MLLM生成既格式正确又空间准确的预测点,我们受眼动点的可变性和多属性画像启发,引入了一致性组相对策略优化(C-GRPO)。在SPA-ADV及其他基准上进行的大量实验证明了本方法的有效性。代码和数据集可在https://github.com/mininglamp-MLLM/PRE-MAP获取。
英文摘要:
Visual selective attention, driven by individual preferences, regulates human prioritization of visual stimuli by bridging subjective cognitive mechanisms with objective visual elements, thereby steering the semantic interpretation and hierarchical processing of dynamic visual scenes. However, existing models and datasets predominantly neglect the influence of subjective cognitive diversity on fixation behavior. Conventional saliency prediction models, typically employing segmentation approaches, rely on low-resolution imagery to generate saliency heatmaps, subsequently upscaled to native resolutions, which limiting their capacity to capture personalized attention patterns. Furthermore, MLLMs are constrained by factors such as hallucinations, making it very costly to strictly adhere to the expected format in tasks involving multiple point predictions, and achieving precise point positioning is challenging. To address these limitations, we present Subjective Personalized Attention for Advertisement Videos, namely SPA-ADV, a large-scale multimodal dataset capturing gaze behaviors from over 4,500 participants varying in age and gender with 486 videos. Furthermore, we propose PRE-MAP, a novel eye-tracking saliency model that characterizes Personalized visual disparities through Reinforcement learning-optimized Eye-tracking, built upon MLLMs and guided by Multi-Attribute user profiles to predict Points. To ensure MLLMs produce prediction points that are both format-correct and spatially accurate, we introduce Consistency Group Relative Policy Optimization (C-GRPO), inspired by the variability in eye movement points and Multi-Attribute profiles. Extensive experiments on SPA-ADV and other benchmarks demonstrate the effectiveness of our approach. The code and dataset are available at \href{https://github.com/mininglamp-MLLM/PRE-MAP}{this URL}.