发表机构
The National High School(国立高中)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对矢量草图模型问题,提出SketchMamba单一因果序列模型,能连续分类草图并生成后续内容。通过密集逐步分类损失实现,在数据集上评估效果良好,参数主干表现出色,消融实验证实密集监督驱动早期预测,统一了识别与生成。
AI 中文摘要
现有矢量草图模型将识别和生成视为单独任务,这给需要实时理解绘图的流接口留下了空白。我们提出了SketchMamba,这是一个单一因果序列模型,它能从任何部分前缀连续分类草图,同时生成其后续内容。通过对选择性状态空间主干应用密集的逐步分类损失来实现。在Quick, Draw!数据集的58类子集上评估,SketchMamba最终步骤准确率达94.93%,渐进准确率曲线下面积(AUC)为0.706,在绘制70%笔触时达到最终准确率的90%。在匹配预算比较中,155万参数主干与因果Transformer相当,优于循环和卷积基线。消融实验证实密集监督机制而非架构本身驱动早期预测能力。结果表明单个因果隐藏状态可统一渐进识别和自回归生成,无需辅助编码器或特定任务分支。
英文摘要
Existing vector-sketch models treat recognition and generation as separate tasks, leaving a gap for streaming interfaces that must understand a drawing as it is being made. We present SketchMamba, a single causal sequence model that continuously classifies a sketch from any partial prefix while simultaneously generating its continuation. We achieve this by applying a dense per-step classification loss to a selective state-space backbone. Evaluated on a 58-class subset of the Quick, Draw! dataset, SketchMamba yields 94.93% final-step accuracy and a progressive-accuracy Area Under the Curve (AUC) of 0.706, crossing 90% of its final accuracy by the time 70% of the strokes are drawn. In a matched-budget comparison, the 1.55 million-parameter backbone ties a causal Transformer while outperforming recurrent and convolutional baselines. Ablations confirm that the dense supervision regime, rather than the architecture alone, drives the early-prediction capability. The results demonstrate that a single causal hidden state can unify progressive recognition and autoregressive generation without auxiliary encoders or task-specific branching.
Comments12 pages, 4 tables, 4 figures