arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34562cs.LG

单层MeMo作为随机汉明核分类器

Single-Layer MeMo as a Randomized Hamming-Kernel Classifier

Alessandro Straziota

首次发表
浏览论文内容

中文总结 AI 辅助

本文证明单层MeMo的理想检索等价于基于位置汉明核的多类分类器,其得分是双重随机草图,并给出误差界与间隔保证,在WikiText-2上展示其准确率、内存与吞吐量的权衡优势。

中文摘要 AI 辅助

MeMo(Zanzotto等人,2025)是一种最近提出的语言模型架构,它将令牌上下文与下一个令牌之间的关联存储在相关矩阵记忆中。在本工作中,我们研究其单层形式,并证明其理想检索规则是基于位置汉明核的多类分类器。MeMo架构使用高斯随机码表示序列特征和输出标签,因此其得分是理想分类器的双重随机草图。在独立的输入和输出码本条件下,我们限定了上下文草图和输出解码引入的误差,刻画了它们对模型和数据参数的依赖,并给出了恢复理想预测的基于间隔的保证。受控模拟支持了分析所预测的趋势。在受限的WikiText-2下一个令牌任务上,我们将单层MeMo与经典基线进行比较,并表明它可以在预测准确性、内存和吞吐量之间提供有用的权衡,特别是在GPU上,其矩阵操作可以并行化。

英文摘要

MeMo (Zanzotto et al., 2025) is a recent language-model architecture that stores associations between token contexts and next tokens in a correlation matrix memory. In this work, we study its single-layer form and show that its ideal retrieval rule is a multiclass classifier based on the positional Hamming kernel. The MeMo architecture represents both the sequence features and the output labels with Gaussian random codes. Its score is therefore a doubly randomized sketch of the ideal classifier. Under independent input and output codebooks, we bound the errors introduced by context sketching and output decoding, characterize their dependence on model and data parameters, and give a margin-based guarantee for recovering the ideal prediction. Controlled simulations support the trends predicted by the analysis. On a restricted WikiText-2 next-token task, we compare single-layer MeMo with classical baselines and show that it can offer a useful trade-off among predictive accuracy, memory, and throughput, particularly on a GPU, where its matrix operations can be parallelized.

发表机构

  • University of Rome “Tor Vergata”(罗马第二大学)

机构由 AI 辅助整理,请以论文原文为准。

↑