标题瓶颈模型
Caption Bottleneck Models
- Dept. of Computer Eng., Middle East Technical University (METU)(中东技术大学计算机工程系)
- Robotics & AI Center (ROMER), METU(中东技术大学机器人与人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出标题瓶颈模型(CaBM),用自由形式自然语言替代固定概念集,通过LMM生成标题并严格基于文本训练分类器,实现无信息泄露的架构并自主发现高质量概念。
AI中文摘要:
概念瓶颈模型(CBM)通过将预测路由到人类可理解的概念层来提供可解释性。然而,为特定数据集定义最优概念集仍然是一个开放的挑战。现有方法依赖于昂贵的专家标注或仅基于类别名称的LLM生成列表。即使是“开放词汇”变体也通常依赖于静态概念集,这限制了发现并引入了标签偏差。此外,传统的CBM常常遭受信息泄露,即未建模的视觉特征绕过瓶颈并损害解释的完整性。为了克服这些限制,我们提出了标题瓶颈模型(CaBM),该框架通过用自由形式的自然语言替代刚性概念层来规避对预定义概念集的需求。通过使用LMM生成的标题表示图像,并严格在此文本上训练分类器,CaBM通过构造确保了无泄露的架构。此外,通过分析训练后的文本分类器,CaBM自主发现高质量的、数据集特定的概念。我们在细粒度和粗粒度基准上的结果表明,CaBM在保持可解释性的同时实现了有竞争力的准确性,而无需外部字典或手动标注的约束。
英文摘要:
Concept Bottleneck Models (CBMs) provide interpretability by routing predictions through a layer of human-understandable concepts. However, defining an optimal concept set for a specific dataset remains an open challenge. Existing approaches rely on expensive expert annotations or LLM-generated lists based solely on class names. Even "open-vocabulary" variants typically depend on static concept sets, which restrict discovery and introduce label bias. Furthermore, traditional CBMs often suffer from information leakage, where unmodeled visual features bypass the bottleneck and compromise the integrity of the explanations. To overcome these limitations, we propose Caption Bottleneck Models (CaBM), a framework that circumvents the need for predefined concept sets by replacing rigid concept layers with free-form natural language. By representing images via LMM-generated captions and training a classifier strictly on this text, CaBM ensures a leakage-free architecture by construction. Additionally, by analyzing the text classifier post-training, CaBM autonomously discovers high-quality, dataset-specific concepts. Our results across fine- and coarse-grained benchmarks demonstrate that CaBM achieves competitive accuracy while preserving interpretability without the constraints of external dictionaries or manual labeling.