AI 中文总结
本文提出一种轻量级CNN,用于实时识别四种手绘几何图形,在2,000张图像数据集上达到96.01%验证准确率,并公开代码与数据。
AI 中文摘要
识别手绘几何图形是草图识别的一个基础子问题,在教育、人机交互和图表数字化等领域具有应用价值。本文介绍了一款桌面应用的设计、实现与评估,该应用使用紧凑的卷积神经网络(CNN)识别四种基本手绘几何图形:圆形、正方形、矩形和三角形。我们独立收集并公开了一个包含2,000张标注的28x28像素形状图像的数据集。分类器由三个卷积块(分别有16、32和64个滤波器)组成,每个卷积块后接最大池化,并包含一个模型内的数据增强阶段(随机水平翻转、旋转和缩放)、一个带有dropout正则化的128单元全连接层,以及一个4类线性输出层,总可训练参数为97,956个。网络使用Adam优化器训练,损失函数为直接基于logits计算的稀疏分类交叉熵。在80/20的训练-验证划分上,模型达到了94.80%的训练准确率和96.01%的验证准确率,验证损失为0.1437。基于Tkinter的图形界面允许用户用鼠标绘制形状,并立即获得带有置信度分数的类别预测。我们将该系统置于更广泛的草图和形状识别文献中,将其准确率与相关的手绘形状分类研究进行比较,并讨论了小型、单一贡献者数据集固有的局限性。完整的源代码、训练好的模型和每类数据集均已公开,以支持可复现性。
英文摘要
Recognizing hand-drawn geometric shapes is a foundational sub-problem of sketch recognition, with applications in education, human-computer interaction, and diagram digitization. This paper presents the design, implementation, and evaluation of a desktop application that recognizes four basic hand-drawn geometric shapes, circle, square, rectangle, and triangle using a compact Convolutional Neural Network (CNN). A dataset of 2,000 labeled 28x28-pixel shape images was collected independently and released publicly. The classifier consists of three convolutional blocks (16, 32, and 64 filters) with max-pooling, an in-model data-augmentation stage (random horizontal flip, rotation, and zoom), a dropout-regularized dense layer of 128 units, and a 4-way linear output layer, totaling 97{,}956 trainable parameters. The network is trained with the Adam optimizer on a sparse categorical cross-entropy objective computed directly on logits. On an 80/20 train-validation split, the model achieves 94.80% training accuracy and 96.01% validation accuracy with a validation loss of 0.1437. A Tkinter-based graphical interface allows a user to draw a shape with the mouse and receive an immediate class prediction with a confidence score. We situate this system within the broader sketch and shape-recognition literature, compare its accuracy against related hand-drawn shape classification studies, and discuss the limitations inherent to a small, single-contributor dataset. The complete source code, trained model, and per-class datasets are released publicly to support reproducibility.