发表机构
University of Sydney(悉尼大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对高级深度学习库抽象神经网络内部机制的问题,本文从零实现自包含神经网络框架,涵盖关键组件。该框架兼具教学功能,应用于多类分类任务性能强大,其设计和模块化使其成为教育及研究的可靠基线。
AI 中文摘要
高级深度学习库的广泛应用虽加速了模型开发,但使神经网络内部机制愈发抽象,导致实际应用与基本理解间产生差距。本文提出一个完全从零实现的自包含神经网络框架,不依赖自动微分或预建深度学习模块。其实现涵盖所有关键组件,包括多层架构、各种激活函数、正则化技术和先进优化器。该框架不仅是揭开前向/反向传播、梯度动态和优化格局神秘面纱的教学工具,应用于多类分类任务时也展现出强大性能,验证了其正确性、数值稳定性及跨不同配置的泛化能力。其可扩展设计和清晰模块化使其成为教育目的和未来研究探索的可靠基线。
英文摘要
The widespread adoption of high-level deep learning libraries, while accelerating model development, has increasingly abstracted away the internal mechanics of neural networks, creating a gap between practical usage and fundamental understanding. To address this, the paper presents a self-contained neural network framework implemented entirely from scratch without relying on automatic differentiation or pre-built deep learning modules. The implementation encompasses all essential components, including multi-layer architectures, diverse activation functions, regularization techniques, and state-of-the-art optimizers. Beyond serving as a pedagogical instrument that demystifies forward/backward propagation, gradient dynamics, and optimization landscapes, the framework demonstrates robust performance when applied to a multi-class classification task, successfully validating its correctness, numerical stability, and generalization across varied configurations. The extensible design and clean modularity further position it as a reliable baseline for educational purposes and future research exploration.