arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2211.11158cs.CVcs.CL

Language in a Bottle:用于可解释图像分类的语言模型引导概念瓶颈

Language in a Bottle: Language Model Guided Concept Bottlenecks for Interpretable Image Classification

  • University of Pennsylvania(宾夕法尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

Yue Yang, Artemis Panagopoulou, Shenghao Zhou, Daniel Jin, Chris Callison-Burch, Mark Yatskar

更新

AI总结:

本文提出LaBo,利用GPT-3生成候选概念、CLIP对齐图像并以次模效用选择瓶颈,在11个数据集上实现可解释且性能可比黑盒的少样本图像分类。

AI中文摘要:

概念瓶颈模型(Concept Bottleneck Models,CBM)是一类内在可解释模型,将模型决策分解为人类可读的概念。它们使人们能够轻松理解模型失败的原因,这是高风险应用中的一项关键特性。CBM需要人工指定概念,并且性能通常不如黑盒对应模型,从而阻碍了其广泛采用。我们解决了这些不足,并首次展示如何在无需人工指定的情况下构建高性能CBM,使其准确率与黑盒模型相近。我们的方法Language Guided Bottlenecks(LaBo,语言引导瓶颈)利用语言模型GPT-3定义一个包含可能瓶颈的巨大空间。给定一个问题领域,LaBo使用GPT-3生成有关类别的事实性句子,以形成候选概念。LaBo通过一种新颖的次模效用高效搜索可能的瓶颈,该效用促进对具有判别性且多样化信息的选择。最终,GPT-3的句子概念可以使用CLIP与图像对齐,从而形成瓶颈层。实验表明,LaBo是视觉识别中重要概念的一种高效先验。在对11个多样化数据集的评估中,LaBo瓶颈在少样本分类方面表现出色:在1 shot设置下,它们的准确率比黑盒线性探针高11.7%,并且在数据更多时性能相当。总体而言,LaBo证明了内在可解释模型可以被广泛应用,并取得与黑盒方法相近甚至更好的性能。

英文摘要:

Concept Bottleneck Models (CBM) are inherently interpretable models that factor model decisions into human-readable concepts. They allow people to easily understand why a model is failing, a critical feature for high-stakes applications. CBMs require manually specified concepts and often under-perform their black box counterparts, preventing their broad adoption. We address these shortcomings and are first to show how to construct high-performance CBMs without manual specification of similar accuracy to black box models. Our approach, Language Guided Bottlenecks (LaBo), leverages a language model, GPT-3, to define a large space of possible bottlenecks. Given a problem domain, LaBo uses GPT-3 to produce factual sentences about categories to form candidate concepts. LaBo efficiently searches possible bottlenecks through a novel submodular utility that promotes the selection of discriminative and diverse information. Ultimately, GPT-3's sentential concepts can be aligned to images using CLIP, to form a bottleneck layer. Experiments demonstrate that LaBo is a highly effective prior for concepts important to visual recognition. In the evaluation with 11 diverse datasets, LaBo bottlenecks excel at few-shot classification: they are 11.7% more accurate than black box linear probes at 1 shot and comparable with more data. Overall, LaBo demonstrates that inherently interpretable models can be widely applied at similar, or better, performance than black box approaches.

补充信息

↑