发表机构
Université de Neuchâtel; Universität Bern; IBM Research; Delft University of Technology(纳沙泰尔大学; 伯尔尼大学; IBM研究院; 代尔夫特理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对自然语言自编码器激活解释中的虚构与写作缺陷问题,提出Flow-NLA方法,通过建模激活分布并利用扩散似然界训练,在保持效用增益的同时减少虚构,提升解释质量。
AI 中文摘要
自然语言自编码器(NLAs)以无监督方式生成模型激活的文本解释:一个言语化器(verbalizer)描述激活,而一个重建器(reconstructor)学习从该文本中恢复激活。在既定的点重建NLA训练方案下,解释在预测模型行为方面变得更加有用,同时也日益引入无根据的细节并表现出写作缺陷。为了分别评估这些变化,我们引入了一个用于非结构化NLA解释的标准化评估框架,测量从解释中可恢复的信息、其主张的上下文支持以及写作质量。为解决虚构和写作缺陷,我们超越了预测单一激活的范畴:解释可以区分可能的激活分布,即使它们的均值和最优点重建奖励相同。我们提出了Flow-NLA,它建模与解释兼容的激活分布,并使用扩散似然界训练言语化器。在Qwen、Gemma和Apertus上,这一更丰富的信号保留了点重建的效用增益,同时抑制了虚构和写作缺陷的增长,开辟了一条改进基于激活的训练方向,以鼓励更具信息性、有依据且可读的解释。代码和评估提示将在接受后公开提供。
英文摘要
Natural Language Autoencoders (NLAs) produce unsupervised text explanations of a model's activations: a verbalizer describes an activation and a reconstructor learns to recover it from this text. Under the established point-reconstruction NLA training recipe, explanations become more useful for predicting model behavior while also increasingly introducing unsupported details and exhibiting writing defects. To assess these changes separately, we introduce a standardized evaluation framework for unstructured NLA explanations, measuring information recoverable from explanations, contextual support for their claims, and writing quality. To address confabulation and writing defects, we move beyond predicting a single activation: explanations can distinguish distributions of possible activations even when their means and optimal point-reconstruction rewards are identical. We introduce Flow-NLA, which models the distribution of activations compatible with an explanation and trains the verbalizer using a diffusion likelihood bound. Across Qwen, Gemma, and Apertus, this richer signal retains the utility gains of point reconstruction while curbing the growth of confabulation and writing defects, opening up a direction for improving activation-derived training to encourage more informative, supported, and readable explanations. Code and evaluation prompts will be made publicly available upon acceptance.
Comments18 pages, 9 figures