AI 中文总结
本文研究ANN符号模式涌现的数学原理,证明其推理逻辑可表述为稀疏符号交互,相关准则使符号模式涌现,为ANN符号解释奠基,还凸显交际学习潜力。
AI 中文摘要
人工神经网络(ANNs)常被视为黑箱模型,可解释性是深度学习的核心挑战,已有诸多工程方法从特征归因、可视化等不同视角对ANN进行近似解释,但长期存在一个开放性问题:ANN的复杂推理逻辑能否被穷尽且简洁地解释为稀疏符号模式。这引出更深层的疑问:符号模式的涌现是反映自然规律还是偶然现象?本文表明,在各类任务上训练的广泛ANN中,其推理逻辑确实可被重新表述为稀疏符号交互;进一步证明,跨任务隐含要求的两个常见数学准则会导致此类稀疏符号交互的涌现,实验证据证实,在不同模型的多数输入样本中这两个准则成立;此外,这些交互的忠实性还通过其强样本间、模型间可迁移性,以及解释ANN整体泛化能力的能力得到验证。本文的理论分析与大量实验为ANN的符号解释提供了坚实基础,为ANN的泛化能力提供了新见解,研究结果还凸显了交际学习(一种可在符号模式层面直接检查和调整ANN推理逻辑的范式,可补充传统端到端学习范式)的潜力;最后,ANN中符号模式的涌现表明,在特定条件下,其他类型黑箱系统中也可能涌现类似符号表示,因为本文的证明不依赖任何特定ANN架构。
英文摘要
Artificial Neural networks (ANNs) are often treated as black-box models, making explainability a central challenge in deep learning. Many engineering methods have been proposed to approximately explain the ANN from various perspectives, such as feature attribution and visualization. However, it remains a long-standing open question whether the complex inference logic of an ANN can be explained exhaustively and concisely as sparse symbolic patterns. This raises a deeper inquiry: does the emergence of symbolic patterns reflect a natural law rather than chance? Here, we show that across a broad class of ANNs trained on diverse tasks, their inference logic can indeed be reformulated as sparse symbolic interactions. We further prove that two common mathematical criteria, which are implicitly required across tasks, lead to the emergence of such sparse symbolic interactions. Empirical evidence confirms that the two criteria hold for the majority of input samples in diverse models. Furthermore, the faithfulness of these interactions is also demonstrated by their strong sample-to-sample and model-to-model transferability, as well as their ability to explain the overall generalization power of ANNs. Our theoretical analysis and extensive experiments provide a solid foundation for symbolic explanations of ANNs, and offer novel insights into the ANN's generalization power. Our findings also highlight the potential of communicative learning, a paradigm in which the inference logic of an ANN can be directly inspected and tuned at the level of symbolic patterns, thus complementing traditional end-to-end learning paradigm. Finally, the observed emergence of symbolic patterns in ANNs suggests that similar symbolic representations may also emerge in other types of black-box systems under certain conditions, because our proof does not depend on any specific ANN architecture.