发表机构
NAVER Cloud; KAIST AI(NAVER云; 韩国科学技术院人工智能)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究构建基准VSysBench测试多模态大语言模型在系统消息下的合规性,发现施加系统消息会降低任务准确率,开放权重模型易受用户冲突影响,视觉约束最难。
AI 中文摘要
多模态大语言模型(Multimodal Large Language Models, MLLMs)的生产部署日益依赖系统消息来管控模型行为。然而现有基准要么仅在文本中评估约束,要么将约束嵌入用户轮次,导致多模态语境下的系统消息依从性在很大程度上未被测量;此外,现有基准还未明确合规性是否以牺牲基础视觉-语言能力为代价。我们推出VSysBench,一个基于MMVet-v2构建的基准,将约束分为5个主要类别和22个子类别,涵盖视觉语境下的文本指令到完全基于视觉的指令,每个类别都配有对应的不对齐配对项以测试指令层级。VSysBench通过联合满意度(Joint Satisfaction Rate, JSR)和跨约束敏感性(Cross-Constraint Sensitivity, CCS)两个维度对每个响应进行联合评分。在16个MLLMs上,我们发现施加系统消息会显著降低基础任务准确率,开放权重模型在用户冲突下合规性崩溃,而顶级专有模型则保持稳定,且基于视觉的约束是所有模型最难的类别。
英文摘要
Production deployments of Multimodal Large Language Models (MLLMs) increasingly rely on system messages to govern model behavior. Yet existing benchmarks either evaluate constraints in text only or embed them into the user turn, leaving system-message adherence in multimodal contexts largely unmeasured; they also leave open whether compliance comes at the cost of foundational vision-language capabilities. We introduce VSysBench, a benchmark built on MMVet-v2 that organizes constraints into 5 main categories and 22 sub-categories, ranging from textual directives in visual contexts to fully vision-grounded ones, each paired with a misaligned counterpart that stress-tests the instructional hierarchy. VSysBench scores each response jointly along two axes, constraint compliance and answer correctness, via the Joint Satisfaction Rate (JSR) and Cross-Constraint Sensitivity (CCS). Across 16 MLLMs, we find that imposing system messages substantially erodes base task accuracy, that compliance collapses under user conflict for open-weight models while remaining stable for top proprietary ones, and that vision-grounded constraints are the hardest category for every model.