arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18096cs.CLcs.LG

MAVEN:基于紧凑对齐评估器的多模态内容宏观社会价值评估框架

MAVEN: A Macro-Societal Value Evaluation Framework of Multimodal Content with Compact Aligned Evaluators

发表机构中国科学技术大学
查看机构详情
  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Zijuan Zhao, Zheren Fu, Hou Xia, Licheng Zhang, Yi Liu, Zhendong Mao

首次发表
浏览论文内容

中文总结 AI 辅助

针对多模态内容宏观社会价值评估难题,提出MAVEN分层框架,构建相关基准与指标,优化评估器,实验表明其2B评估器表现接近前沿闭源VLMs,提供了可扩展评估路径。

中文摘要 AI 辅助

评估多模态内容是否契合和平、正义、自由等宏观社会价值已成为日益紧迫的挑战。现有框架大多局限于安全导向的分类体系、仅文本的心理测量探针或单标签分类。因此,我们提出MAVEN,这是一个基于国际人权文书和文化价值理论的多模态内容宏观社会价值评估分层框架。MAVEN将价值划分为6个主要维度和72个二级指标,支持多级定量评分。基于MAVEN,我们构建了一个经人工验证的多模态基准和一个软匹配指标,用于评估VLMs在各价值维度的表现。针对评估器优化,我们提出了一种适用于评估器蒸馏的多级别偏好优化的跨度自适应变体,以及推理时的无训练多角色共识策略。我们在基准上评估了现有开源和闭源VLMs,揭示了它们在宏观社会价值判断中的共同倾向和明显差异。实验表明,我们的2B紧凑评估器与同系列8B评估器表现相当,且接近前沿闭源VLMs,为可扩展的宏观社会价值评估提供了可行路径。我们的SA-MDPO实现和MacroValue-Bench可在该httpsURL获取。

英文摘要

Assessing whether multimodal content aligns with macro-societal values, such as peace, justice, and freedom, has become an increasingly urgent challenge. Existing frameworks are largely confined to safety-oriented taxonomies, text-only psychometric probes, or single-label classification. Therefore, we propose MAVEN, a hierarchical framework for macro-societal value evaluation of multimodal content, grounded in international human-rights instruments and cultural value theory. MAVEN organizes values into 6 primary dimensions and 72 secondary indicators, supporting multi-level quantitative scoring. Building on MAVEN, we construct a human-verified multimodal benchmark and a soft-match metric to evaluate VLMs' assessments across value dimensions. For evaluator optimization, we propose a span-adaptive variant of multi-level preference optimization for evaluator distillation, together with a training-free multi-role consensus strategy at inference time. We evaluate existing open- and closed-source VLMs on our benchmark, revealing shared tendencies and clear differences in macro-societal value judgments. Experiments show that our compact 2B evaluator matches its 8B counterpart in the same family and approaches frontier closed-source VLMs, offering a practical path toward scalable macro-societal value evaluation. Our SA-MDPO implementation and MacroValue-Bench are available at https://github.com/zzzzzzzzjj/MAVEN.

补充信息

↑