arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07594cs.CLcs.AI

扩展固有可解释的语言模型

Scaling Inherently Interpretable Language Models

Guide Labs Team, Andreas Madsen, Aya Abdelsalam Ismail, Giang Nguyen, Isaac Plant, Muawiz Chaudhary, Nathaniel Monson, Saqib Azim, Zhichen Guo, Julius Adebayo

首次发表
浏览论文内容

中文总结 AI 辅助

本研究将可解释性作为训练约束而非事后补充,在多规模计算下验证了其随模型能力同步扩展的特性,通过Steerling-8B模型实现闭环干预,且性能优于更大计算量的同类模型。

中文摘要 AI 辅助

可解释性常被视为对模型能力的一种“负担”:语言模型被训练为不透明系统,之后再通过可靠性难以验证的方法进行事后解释。本研究对这一前提提出挑战,并非对模型进行逆向工程,而是将可解释性作为训练流程的约束条件,与语言建模目标一同优化。在三个数量级的计算规模下,针对自回归模型和扩散语言模型,可解释性随模型能力同步扩展而非相悖;令人惊讶的是,模型表征会随规模扩大变得更解耦,且与人类可理解的概念更对齐。我们用Steerling-8B(一种带有因果注意力掩码的扩散语言模型)实现了该训练方案:对于任意一组生成的token,Steerling-8B会将输出归因于相关输入token、人类可理解的概念及训练数据,这支持闭环干预:通过概念或特征归因诊断输出,检索相似训练数据,并通过概念引导修正行为而无需重新训练。Steerling-8B在计算量比其多2-16倍的公开同类模型中仍具竞争力,这表明了一种不同的扩展范式:可解释性可被融入训练设计,且会随规模提升而优化。

英文摘要

Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, with methods whose reliability is difficult to establish. In this work, we challenge this premise. Rather than reverse-engineering a model, we make interpretability a constraint of the training pipeline, optimized alongside the language modeling objective. Across three orders of magnitude of compute, on both autoregressive and diffusion language models, interpretability scales with capability rather than against it. Surprisingly, model representations become more disentangled and aligned with human-understandable concepts with scale. We instantiate the training-time recipe with Steerling-8B, a diffusion language model with a causal attention mask. For any group of generated tokens, Steerling-8B attributes the output to relevant input tokens, human-understandable concepts, and training data. This enables closed-loop intervention: diagnose an output through its concept or feature attribution, retrieve similar training data, and correct the behavior through concept steering without retraining. Steerling-8B remains competitive with open peer models trained on substantially 2-16x more compute, suggesting a different scaling paradigm: interpretability can be designed into training, and it improves with scale.

发表机构

  • Guide Labs(指南实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑