MM-IFEval-Pro:一个面向视觉语言模型的多语言且抗攻击的指令遵循基准
MM-IFEval-Pro: A Multilingual and Attack-Resistant Benchmark for Instruction-Following in Vision-Language Models
浏览论文内容
中文总结 AI 辅助
针对现有多模态指令遵循基准语言覆盖有限、对抗场景不足的问题,本文提出多语言抗攻击基准MM-IFEval-Pro,构建含中文及对抗指令的强化学习训练集,提升模型性能并实现跨任务跨语言泛化。
中文摘要 AI 辅助
随着视觉语言模型(VLMs)在图像理解、跨模态推理和复杂指令执行方面快速发展,指令遵循能力已成为衡量其可靠性和实用性的关键指标。然而,现有的多模态指令遵循基准仍存在语言覆盖范围有限、对抗安全场景不足的问题,无法满足现实世界多语言和安全敏感场景的评估需求。为解决这些不足,本文提出MM-IFEval-Pro,这是一个多模态指令遵循基准,涵盖中文和英文任务以及多样化的指令劫持案例。MM-IFEval-Pro包含4个主要任务类别、24个子类别,以及8个指令类别、52个子类别,每个样本平均包含3.0个约束条件,以逼真模拟复杂指令场景。我们进一步构建了一个包含中文和对抗指令的强化学习训练集,该训练集显著提升了模型在MM-IFEval-Pro上的性能,且能有效迁移到其他主流多模态基准,展现出强大的跨任务和跨语言泛化能力。
英文摘要
As vision-language models (VLMs) rapidly advance in image understanding, cross-modal reasoning, and complex instruction execution, instruction-following capability has become a key indicator of their reliability and practicality. However, existing multimodal instruction-following benchmarks still suffer from limited language coverage and insufficient adversarial safety scenarios, making them inadequate for evaluating real-world multilingual and safety-sensitive settings. To address these gaps, we present MM-IFEval-Pro, a multimodal instruction-following benchmark covering Chinese and English tasks as well as diverse instruction hijacking cases. MM-IFEval-Pro includes 4 major task categories and 24 subcategories and 8 instruction categories with 52 subcategories, with each sample containing an average of 3.0 constraints to realistically simulate complex instruction scenarios. We further construct a reinforcement-learning training set enriched with Chinese and adversarial instructions, which significantly improves model performance on MM-IFEval-Pro and transfers effectively to other mainstream multimodal benchmarks, demonstrating strong cross-task and cross-language generalization.
发表机构
- Huawei Technologies Ltd.(华为技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。