发表机构
Pohang University of Science and Technology; Upstage AI(浦项科技大学; Upstage AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出PRIMEBench,一种视觉感知的分层基准压缩框架,通过四阶段剪枝在保留模型排名的同时大幅降低VLM评估成本,并揭示评估行为随模型面板演变的规律。
AI 中文摘要
对视觉-语言模型(VLM)的全面评估已变得极其昂贵,因为基准测试覆盖了越来越广泛的能力范围,且新模型以不停歇的速度涌现。保留模型排名的基准压缩方法在语言模型领域已得到充分研究,但对于VLM,这一问题仍探索不足。我们提出了PRIMEBench(多模态评估的冗余项剪枝),一种视觉感知的分层基准压缩框架,在保留模型排名的同时大幅降低评估成本。该分层框架分四个阶段运行:数据清洗,移除无需图像即可回答的项和全对项;类别代表性选择,为每个能力类别挑选一个基准;使用视觉感知方差(VAW)进行项剪枝;以及类别数量剪枝。VAW结合了模型间方差与仅从多模态嵌入计算的视觉依赖分数,同时鼓励每个基准内覆盖多样化的项。在从项选择中留出的模型上,它在发布的5%保留率下具有最高的平均保真度。分层设计使从业者可以在任何阶段停止以匹配其计算预算;发布的套件移除了超过97%的项,同时保留了模型排名。除了压缩之外,我们的分析还展示了VLM评估如何随着模型面板的增长和演变而变化,为设计未来更高效、对模型更替更稳健、并明确评估侧剪枝限制的基准提供指导。
英文摘要
Thorough evaluation of vision-language models (VLMs) has become prohibitively expensive, as benchmarks span an ever-broader spectrum of capabilities and new models arrive at a relentless pace. Benchmark compression methods that preserve model rankings at a fraction of the cost are well studied for language models, but for VLMs the question remains under-explored. We present PRIMEBench (Pruning Redundant Items for Multimodal Evaluation), a vision-aware hierarchical benchmark compression framework that substantially reduces evaluation cost while preserving model rankings. This hierarchical framework operates in four stages: data cleaning to remove items answerable without the image and all-correct items, category representative selection to pick one benchmark per capability category, item pruning with Vision-Aware Variance (VAW), and category-count pruning. VAW combines inter-model variance with a vision-dependence score computed from multimodal embeddings alone, while encouraging coverage of diverse items within each benchmark. On models held out from item selection, it has the highest mean fidelity at the released 5% retention. The hierarchical design lets practitioners stop at any stage to match their compute budget; the released suite removes over 97% of items while preserving model rankings. Beyond compression, our analyses show how VLM evaluation behaves as model panels grow and evolve, providing guidance for designing future benchmarks that are more efficient, robust to model turnover, and explicit about the limits of evaluation-side pruning.
CommentsPreprint