发表机构
Alibaba International Digital Commerce Group(阿里巴巴国际数字商业集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对视觉语言模型的上下文变量高估问题,提出GeoReward奖励模型,通过三种机制缓解该偏差,在跨市场广告偏好预测任务中性能优于现有基线。
AI 中文摘要
视觉语言模型(VLM)在多模态任务中表现出色,但仍存在一种微妙却影响重大的失效模式:它们倾向于高估主导性的视觉-文本线索,同时低估稀疏但对决策至关重要的上下文变量。我们将这一问题称为上下文变量高估(Contextual Variable Overestimation,CVE),在跨不同地理市场预测广告图像偏好等实际应用中尤为明显。例如,当要求VLM在为不同国家定制的两张产品图像中做出选择时,它往往会输出一致的结果,忽略真实的区域差异。这种失效的原因在于,产品属性、密集图像斑块等大量普遍存在的信号,会压倒编码市场特定上下文的少数关键 token。为解决CVE问题,我们首先收集了一个新的多模态数据集,包含真实广告创意及其在多个国家的点击率表现。随后,我们推出了GeoReward,一个用于预测不同地理市场广告图像偏好的奖励模型,它整合了三个专为该任务设计的机制:(1)市场感知检索增强(Market-Aware Retrieval Augmentation);(2)上下文引导视觉调制(Context-Guided Visual Modulation);(3)选择性敏感损失(Selective Sensitivity Loss)。此外,我们展示了GeoReward如何引导视觉语言模型的强化学习(RL)微调,以生成文本到图像模型的背景设计,从而产出具有市场感知的广告创意。实验验证了我们的框架可缓解CVE问题,且性能优于现有基线。本研究不仅诊断了VLM对主导感知特征存在的系统性偏差,还为稀疏上下文变量主导决策的应用提供了针对性解决方案。
英文摘要
Vision-language models excel in many multimodal tasks but remain prone to a subtle yet impactful failure mode: they tend to overestimate dominant visual-textual cues while underestimating sparse but decision-critical contextual variables. This issue, which we term Contextual Variable Overestimation (CVE), becomes particularly evident in real-world applications such as predicting advertisement image preferences across diverse geographic markets. For instance, when a VLM is asked to choose between two product images tailored for different countries, it often defaults to a consistent output, ignoring ground-truth regional variations. This collapse occurs because pervasive high-volume signals, such as product attributes and dense image patches, overwhelm the few but critical tokens that encode market-specific context. To address CVE, we first collect a new multimodal dataset of real advertising creatives and their click-through performance across multiple countries. We then introduce GeoReward, a reward model designed to predict ad image preferences across diverse geographic markets. GeoReward integrates three purpose-built mechanisms: (1) Market-Aware Retrieval Augmentation, (2) Context-Guided Visual Modulation, (3) Selective Sensitivity Loss. Furthermore, we demonstrate how GeoReward can guide the fine-tuning of RL for a VLM to generate background designs for text-to-image models, producing market-aware advertising creatives. Experiments validate that our framework mitigates CVE and outperforms existing baselines. This work not only diagnoses a systematic bias in VLMs toward dominant perceptual features but also delivers a targeted solution for applications where sparse contextual variables govern decision-making.