AI 中文总结
该研究针对多模态指令多要求重要性不等的问题,提出PRISM四阶段数据合成框架,结合PRISM-Eval评估,提升多模态大模型的多规则优先级感知指令遵循能力。
AI 中文摘要
现实世界的多模态指令通常包含多个重要性不等的要求,但大多数多模态训练数据仍将指令遵循简化为回答一个独立的问题。我们通过评分标准理解研究这一差距,该研究将模型不是作为根据评分标准进行评估的生成器,而是作为遵循评分标准的执行器:给定一张图像和一个带类型的、有优先级的评分标准,模型必须验证每条规则,然后才能做出整体判断。为支持该设置,我们提出PRISM,这是一个四阶段的数据合成框架,可生成角色-任务对、前缀引导的规则集、经过质量筛选的评分标准以及结构化验证轨迹。我们进一步引入PRISM-Eval,其宽松和严格指标使用针对固定标签的确定性匹配,因此不需要推理时的评判模型。仅使用10000个合成样本,PRISM将Qwen3-VL-4B在PRISM-Eval上的严格准确率从9.5%提升至30.1%,同时保持在通用基准上的平均性能,且该提升可迁移到四个额外的开源多模态大语言模型(MLLM),涵盖密集型和混合专家(MoE)架构,表明结构化评分标准监督是迈向多规则、感知优先级的多模态指令遵循的可扩展路径。
英文摘要
Real-world multimodal instructions often bundle multiple requirements with unequal importance, yet most multimodal training data still reduce instruction following to answering one self-contained question. We study this gap through rubric comprehension, which casts the model not as a generator measured against rubrics but as an executor that follows them: given an image and a typed, prioritized rubric, the model must verify each rule before producing an overall judgment. To support this setting, we propose PRISM, a four-stage data synthesis framework that produces persona--task pairs, prefix-guided rule sets, quality-filtered rubrics, and structured verification traces. We further introduce PRISM-Eval, whose Loose and Strict metrics use deterministic matching against fixed labels and therefore require no inference-time judge model. With only 10K synthesized samples, PRISM lifts Qwen3-VL-4B from 9.5% to 30.1% Strict accuracy on PRISM-Eval while preserving average performance on general benchmarks, and the gains transfer to four additional open-source MLLMs across dense and MoE architectures, suggesting that structured rubric supervision is a scalable path toward multi-rule, priority-aware multimodal instruction following.