arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CBX-Bench:用于评估概念瓶颈模型解释的人类对齐多模态大语言模型委员会

CBX-Bench: A Human-Aligned MLLM Council for Benchmarking Concept Bottleneck Model Explanations

Yusuf Meric Karadag, Gulay Oklan, Seref Baris Cagliyan, Umut Ozdemir, Emre Akbas

arXiv 2608.15404首次发表:更新:

AI 中文总结

该研究开发了人类对齐的多模态大语言模型(MLLM)委员会,构建了CBX-Bench基准,实现了概念瓶颈模型(CBM)解释的可扩展定量评估,其在人类偏好排名上的恢复率达70%以上。

AI 中文摘要

概念瓶颈模型(CBMs)旨在通过用人类可理解的概念表达预测结果,实现视觉分类的可解释性。尽管可解释性是CBMs的核心动机,但目前仍主要通过下游分类准确率对其进行评估,仅辅以孤立的定性示例,这凸显了对定量评估指标的迫切需求。这一挑战因大规模概念标注的真实值不可行,以及概念列表因缺乏共识而具有开放性而进一步复杂化。为填补这一空白,我们开发了一个多模态大语言模型(MLLM)委员会,该委员会在给定图像及其CBM解释后,会生成解释质量评分。为确立并验证该委员会,我们首先开展了一项人类研究,以建立CBM解释质量的真实值参考:针对每张图像,标注人员比较LF-CBM、VLG-CBM和CBM-Suite中两个模型的解释,选择更有用的解释,或标记为同等好或同等差,最终在CUB-200、ImageNet-100和Places365上的900个图像对比项中获得2700个判断。与人类参考相比,由五个开源权重MLLMs组成的委员会恢复了超过70%的严格人类偏好排名,在人类标注者一致同意的项上这一比例升至83%。基于这一经验证的委员会,我们推出了CBX-Bench,这是一个公共基准和排行榜:新CBM的作者可提交其模型的解释,CBX-Bench将使用该委员会对其进行评分,并维护数据集层面的解释质量排名。因此,CBX-Bench提供了一种超越准确率和孤立定性示例的、人类对齐且可扩展的CBM解释评估方式,该基准可通过指定URL访问。

英文摘要

Concept Bottleneck Models (CBMs) are designed to make visual classification interpretable by expressing predictions through human-understandable concepts. Although interpretability is the central motivation for CBMs, they are still largely evaluated as predictive models by downstream classification accuracy, supplemented by isolated qualitative examples. This highlights a pressing need for quantitative measures, a challenge complicated by the infeasibility of ground-truth concept annotation at scale and the open nature of concept lists due to a lack of consensus. To fill this gap, we develop a multimodal large language model (MLLM) council that, given an image and its CBM explanation, produces an explanation quality score. To ground and validate the council, we first conduct a human study to establish a ground-truth reference for CBM explanation quality: for an image, annotators compare explanations from two of LF-CBM, VLG-CBM, and CBM-Suite and choose the more useful one, or mark them as equally good or equally bad, yielding 2700 judgments over 900 image-comparison items on CUB-200, ImageNet-100, and Places365. Against this human reference, our five-model council, consisting of open-weight MLLMs, recovers over 70% of strict human preference rankings, rising to 83% on items where human annotators unanimously agree. Building on this validated council, we introduce CBX-Bench, a public benchmark and leaderboard: authors of new CBMs can submit their model's explanations, and CBX-Bench scores them with the council and maintains dataset-level rankings of explanation quality. CBX-Bench thus provides a human-aligned, scalable evaluation of CBM explanations beyond accuracy and isolated qualitative examples. The benchmark is available at https://github.com/meric-karadag/cbx-bench.

CommentsAccepted to the Explainable Computer Vision (eXCV) Workshop at ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑