概念瓶颈大型语言模型
Concept Bottleneck Large Language Models
AI总结:
提出概念瓶颈大型语言模型框架,将内在可解释性集成到LLMs中,在文本分类和生成任务上实现可解释推理、精确概念检测与安全控制,增强模型安全性与可信度。
AI中文摘要:
我们引入了概念瓶颈大型语言模型(CB-LLMs),这是一个用于构建本质上可解释的大型语言模型(LLMs)的新框架。与依赖有限的事后解释的传统黑盒LLMs相比,CB-LLMs将内在可解释性直接集成到LLMs中——在提供准确解释的同时,实现可扩展性和透明度。我们为两个关键的NLP任务构建了CB-LLMs:文本分类和文本生成。在文本分类中,CB-LLMs与传统黑盒模型具有竞争力,有时甚至表现更优,同时提供明确且可解释的推理。对于更具挑战性的文本生成任务,CB-LLMs中的可解释神经元能够实现精确的概念检测、受控生成和更安全的输出。嵌入的可解释性使用户能够透明地识别有害内容、引导模型行为并遗忘不需要的概念——显著增强了LLMs的安全性、可靠性和可信度,这些关键能力在现有模型中明显缺失。我们的代码可在https://github.com/Trustworthy-ML-Lab/CB-LLMs获取。
英文摘要:
We introduce Concept Bottleneck Large Language Models (CB-LLMs), a novel framework for building inherently interpretable Large Language Models (LLMs). In contrast to traditional black-box LLMs that rely on limited post-hoc interpretations, CB-LLMs integrate intrinsic interpretability directly into the LLMs -- allowing accurate explanations with scalability and transparency. We build CB-LLMs for two essential NLP tasks: text classification and text generation. In text classification, CB-LLMs is competitive with, and at times outperforms, traditional black-box models while providing explicit and interpretable reasoning. For the more challenging task of text generation, interpretable neurons in CB-LLMs enable precise concept detection, controlled generation, and safer outputs. The embedded interpretability empowers users to transparently identify harmful content, steer model behavior, and unlearn undesired concepts -- significantly enhancing the safety, reliability, and trustworthiness of LLMs, which are critical capabilities notably absent in existing models. Our code is available at https://github.com/Trustworthy-ML-Lab/CB-LLMs.