AI 中文总结
本文提出150万参数的字节级多头内容分类器pico-type,无需分词器等组件,可单次预测7项内容属性,在多基准测试中准确率大幅优于仅合成数据训练的模型,相关成果已开源。
AI 中文摘要
我们推出pico-type,这是一种拥有约150万参数的字节级多头内容分类器,可在单次前向传播中从原始UTF-8字节同时预测7项内容属性。pico-type直接在字节层面运行,无需分词器、子词词汇表或预训练嵌入,可对粗类型(12类)、模态(8类)、子类型(24类)、代码语言(62种)、文本语言(30种)、文件MIME类型(90类)及风险标记(6标签多标签:API密钥、JWT、密码、电子邮件、电话号码、SSH密钥)进行分类。其架构结合了学习到的字节嵌入、3个感受野递增的卷积块、2个带旋转位置编码的双向注意力层,以及为7个套娃式分类头提供输入的统计池化层。4种分层变体(tiny/small/base/pro)共享同一主干,采用16至576维度的切片表示,生成的ONNX导出文件小于210KB,CPU推理耗时在10毫秒以内。pico-type在合成模板与真实数据的混合数据集(8709个GitHub代码样本、5000篇维基百科文章)上训练,在The Heap基准(24种语言)上实现了60.3%的代码语言准确率,在维基百科(30种语言)上实现了98.2%的文本语言准确率,较仅用合成数据训练的基准分别提升了57和79个百分点;基于格式的头(粗类型、模态、子类型、文件MIME、风险)在合成基准上保持了100%的准确率。该模型、代码及预训练权重以Apache 2.0许可发布。
英文摘要
We introduce pico-type, a byte-level multi-head content classifier with approximately 1.5 million parameters that simultaneously predicts seven content properties from raw UTF-8 bytes in a single forward pass. Operating directly at the byte level -- no tokenizer, no subword vocabulary, no pretrained embeddings -- pico-type classifies coarse type (12 classes), modality (8), subtype (24), code language (62), text language (30), file MIME type (90), and risk flags (6-label multi-label: API keys, JWTs, passwords, emails, phone numbers, SSH keys). The architecture combines a learned byte embedding, three convolutional blocks with growing receptive fields, two bidirectional attention layers with rotary position encodings, and a statistical pooling layer feeding seven Matryoshka-style classification heads. Four tiered variants (tiny/small/base/pro) share the same trunk with sliced representations from 16 to 576 dimensions, yielding ONNX exports under 210 KB and CPU inference under 10 ms. Trained on a mixture of synthetic templates and real-world data (8709 GitHub code samples, 5000 Wikipedia articles), pico-type achieves 60.3 percent code language accuracy on The Heap benchmark (24 languages) and 98.2 percent text language accuracy on Wikipedia (30 languages) -- improvements of +57 and +79 percentage points respectively over the synthetic-only baseline. Format-based heads (coarse, modality, subtype, file_mime, risk) maintain 100 percent accuracy on synthetic benchmarks. The model, code, and pretrained weights are released under Apache 2.0.
Comments14 pages, 1 figure, 8 tables