生物学领域AI模型的高效自动可解释性
Efficient Auto-Interpretability of AI Models in Biology
AI总结:
该研究整合潜在特征一致性、可描述性与预测能力的评估流程,在Boltzmann-1 Pairformer上实现高效可解释性,降低评估成本并恢复过半可解释潜在特征,同时发现跨种子稳定性可能偏好结构相关特征。
AI中文摘要:
稀疏自编码器(SAEs)及其他可解释性方法可通过解释生物学等领域AI模型的超人类能力,将这些模型转化为科学发现引擎。然而,只有当潜在特征满足三个条件时才有用:是否具有一致性、是否可描述、该描述是否具有预测能力,这些问题常被混淆。我们将其整合为单一流程,并报告各阶段所需的实用创新:第一,跨种子字典稳定性优先确定哪些潜在特征值得投入资源研究;第二,入侵者检测任务判断潜在特征激活的示例是否具有可识别模式;第三,单独一轮提出候选生物学描述,我们将其转化为可在计算机上验证的可证伪预测。在Boltzmann-1 Pairformer主干模型上部署后,稳定性优先策略每个潜在特征的评估资源减少约4.4倍,测量成本降低5.2倍,同时恢复了超过一半的潜在特征;外部检查显示,发现的基序(motifs)显著富集了其宣称的注释。结果还表明存在潜在矛盾:跨种子稳定性可能更倾向于选择结构相关特征,而非功能相关特征。
英文摘要:
Sparse autoencoders (SAEs), and other interpretability methods could turn AI models in Biology and other fields into engines of scientific discovery by explaining the superhuman capabilities of those models. However, a latent is only useful if we know three things: whether it is coherent, whether it can be described, and whether that description has predictive power. These questions are routinely conflated. We assemble them into a single pipeline and report the practical innovations each stage required. First, cross-seed dictionary stability prioritises which latents are worth spending resources to investigate. Second, an intruder-detection task asks whether a latents activating examples share a recognizable pattern. Third, a separate pass proposes a candidate biological description which we convert into falsifiable predictions which can be tested in silico. Deployed on the Boltz-1 Pairformer trunk, stability prioritisation finds interpretable latents using about 4.4 times fewer latent evaluations each, and at 5.2 times lower measured cost, while recovering over half of them, and the external check shows the surfaced motifs are significantly enriched for their claimed annotations. The results also suggest a possible tension: the cross- seed stability might be selecting for some types of features, like structure-related ones, much more than others, such as function-related features.