arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15740cs.CVcs.AIcs.LGcs.MM

通过隐式文化对齐奖励建模消除文本到图像评估中的偏差

Debiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling

Bo-An Chang, Yu-Chih Chen

首次发表
浏览论文内容

中文总结 AI 辅助

研究文本到图像评估中文化真实性问题,提出基于轻量级多模态大语言模型构建的隐式文化对齐奖励模型,集成隐式文化探测器与跳跃连接交叉注意力机制,实验证明该模型准确率高、速度快,能为偏好优化管道提供有效信号。

中文摘要 AI 辅助

随着文本到图像(T2I)系统的迅速发展,评估合成内容的文化真实性对于公平且可信的生成式人工智能变得越发重要。现有T2I评估指标和多模态评判通常依赖视觉语义表示,这不足以体现隐式文化规范,导致偏好判断有偏差且遗漏细粒度文化线索。此外,基于视觉问答(VQA)的评估器通常依赖自回归文本生成,限制了实时奖励建模的扩展性。为解决这些局限,我们引入基于轻量级42亿参数多模态大语言模型(MLLM)构建的隐式文化对齐奖励模型。我们的框架将隐式文化探测器与跳跃连接交叉注意力(SkipCA)机制集成,使后期语义特征能直接关注早期视觉表示,更好地保留文化显著细节。对来自CulturalFrames基准的3323对具有挑战性且精心挑选的图像对进行评估表明,我们的方法实现了80.54%的成对准确率,皮尔逊和肯德尔相关系数分别为0.546和0.377,优于代表性视觉语言指标和基于MLLM的评估器。此外,通过绕过自回归文本生成,我们的模型在本地推理设置下以0.21秒处理每次评估,比基于标准VQA的评估器快10倍。这些结果表明,所提出的奖励模型可为诸如从人类反馈强化学习和直接偏好优化等偏好优化管道提供高效且具有文化意识的标量信号。

英文摘要

As Text-to-Image (T2I) systems rapidly advance, evaluating the cultural authenticity of synthesized content has become increasingly important for fair and trustworthy generative AI. Existing T2I evaluation metrics and multimodal judges often rely on visual-semantic representations that underrepresent implicit cultural norms, leading to biased preference judgments and the omission of fine-grained cultural cues. In addition, visual question answering (VQA)-based evaluators typically depend on autoregressive text generation, which limits their scalability for real-time reward modeling. To address these limitations, we introduce an Implicit Cultural Alignment Reward Model built upon a lightweight 4.2-billion-parameter Multimodal Large Language Model (MLLM). Our framework integrates an Implicit Cultural Probe with a Skip-connection Cross-Attention (SkipCA) mechanism, enabling late-stage semantic features to directly attend to early-stage visual representations and better preserve culturally salient details. Evaluations on 3,323 challenging and carefully curated image pairs from the CulturalFrames benchmark show that our approach achieves 83.49% pairwise accuracy, with Pearson and Kendall correlation coefficients of 0.5268 and 0.3749, respectively, outperforming representative vision-language metrics and MLLM-based evaluators. Moreover, by bypassing autoregressive text generation, our model processes each evaluation in 0.21 seconds under our local inference setup, achieving a $10\times$ speedup over standard VQA-based evaluators. These results suggest that the proposed reward model can provide an efficient and culturally aware scalar signal for preference optimization pipelines such as Reinforcement Learning from Human Feedback and Direct Preference Optimization. Additional resources are available on our project page at https://bensonch1214.github.io/Implicit_Cultural_Alignment/.

发表机构

  • National Tsing Hua University(国立清华大学)
  • National Yang Ming Chiao Tung University(国立阳明交通大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑