发表机构
Frederick University(弗雷德里克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出无注意力的混响模型,在字节级语言建模上优于参数匹配的Transformer,引入FORM DISTANCE衡量文本真实性,发现解码策略对生成性能的影响大于架构,且增益依赖训练语料库内的短语。
AI 中文摘要
Kathleen系列第1、2篇论文表明,由波形表编码器和多尺度混响状态构建的字节级无注意力架构,在约450-700K参数、无需预训练的情况下,可在分类任务上媲美强基线模型。本文探究相同组件能否用于生成任务:(1)缩放:在字节级语言建模(WikiText-103、原始UTF-8、无分词器)任务中,混响模型在所有测试的数据集规模(2-512 MB)下均优于参数匹配的Transformer,例如在512 MB数据、约0.5M参数时,其比特每字节值为1.84,而Transformer为2.04;Transformer需要超过512 MB数据才能达到该无注意力模型从32 MB数据中学到的效果。(2)测量:我们引入FORM DISTANCE,一种用于衡量“是否像文本”的非参数、抗作弊工具:以人类文本的九个统计轴定义参考云,五种构造的假文本均被成功拒绝。(3)生成:解码策略主导架构性能——扩大采样器可使同一模型的距离减半(从3.17降至1.52),检索增强解码方案无需训练步骤即可让冻结模型进一步提升(从1.52降至1.14);消融实验表明增益源于稀疏短语剂量本身,而非选择门。该增益存在明确边界条件:短语必须来自模型自身的训练语料库——40倍大的外部语料库毫无帮助,该效应在注意力孪生模型中也存在,符合上下文集成是规模相关能力的结论。我们还报告了四项未起作用的架构改进,以及一个参数量仅为学习表五分之一、达到其Top-1准确率94%的计算词典。所有实验均离线运行,可在免费Kaggle T4上复现。
英文摘要
Papers 1-2 of the Kathleen series showed that a byte-level, attention-free architecture built from a wavetable encoder and multi-scale reverberant state can match strong baselines on classification at ~450-700K parameters, without pretraining. We ask whether the same ingredients can generate. (1) Scaling: on byte-level language modeling (WikiText-103, raw UTF-8, no tokenizer), the reverberant model beats a parameter-matched transformer at every dataset scale measured (2-512 MB), e.g. 1.84 vs 2.04 bits/byte at 512 MB with ~0.5M parameters; the transformer needs more than 512 MB to match what the attention-free model learns from 32 MB. (2) Measurement: we introduce FORM DISTANCE, a non-parametric, gaming-resistant instrument for "reads like text": nine statistical axes of human text define a reference cloud, and five constructed fakes are all rejected. (3) Generation: decoding policy dominates architecture -- widening the sampler halves the same model's distance (3.17 to 1.52), and a retrieval-augmented decoding scheme takes the frozen model further (1.52 to 1.14) with no training step involved; the ablation attributes the gain to the sparse phrase dose itself, not the selection gate. The gain has a sharp boundary condition: the phrases must come from the model's own training corpus -- a 40x larger foreign library helps not at all, an effect the attention twin shares, consistent with in-context integration being a capability of scale. We also report four architectural additions that did not help, and a computed lexicon reaching 94% of a learned table's top-1 accuracy at one fifth of the parameters. Everything runs offline; all experiments are reproducible on a free Kaggle T4.
CommentsPaper 3 of the Kathleen series. 11 pages, 3 figures. All experiments reproducible on a free Kaggle T4