arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TinyCeNN-LM:基于CeNN启发的细胞递归层对预训练注意力的质量门控转换

TinyCeNN-LM: Quality-Gated Conversion of Pretrained Attention with CeNN-Inspired Cellular-Recurrent Layers

Kabeh Mohsenzadegan, Vahid Tavakkoli, Kyandoghere Kyamakya

arXiv 2609.21139首次发表:更新:

发表机构

University of Klagenfurt; Institute for Smart System Technologies; Faculte Polytechnique Universite de Kinshasa(克拉根福大学; 智能系统技术研究所; 金沙萨大学理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出TinyCeNN-LM框架,通过质量门控将预训练模型注意力转换为CeNN启发细胞递归层,在保持性能的同时减少缓存,支持保守结构转换。

AI 中文摘要

在预训练语言模型中替换注意力机制是一个兼容性问题:一个看似合理的替代方案可能会改变后续层所期望的表示。TinyCeNN-LM引入了一个基于CeNN启发的细胞递归层的质量门控后训练转换框架,该框架具有有界局部处理、紧凑递归记忆、路由、融合以及接受或回滚验证。研究了三种实现:集成记忆(Integrated Memory)、记忆融合(MemoryFusion)和PDelta3-GDN2-CLVR+Local32。严格的PDelta3转换仅在表示和NLL标准通过固定阈值时接受某一层。在SmolLM2-135M上,第0-2层被接受,累计ΔNLL=+0.01209,而第3层尽管NLL可接受但因表示保真度失败而被拒绝。在Qwen3.5-0.8B上,全注意力层3、7和11被接受,最终ΔNLL=+0.02073。集成记忆将困惑度保持在-0.07%至+0.93%范围内,同时将总缓存减少最多6.01%。对转换后的Qwen版本进行200项采样下游合理性检查,总体准确率为28.5%至32.0%。结果支持保守的、质量门控的结构转换,而非普遍的注意力替换或加速。

英文摘要

Replacing attention in a pretrained language model is a compatibility problem: a plausible substitute may alter representations expected by later layers. TinyCeNN-LM introduces a \emph{quality-gated post-training conversion} framework using CeNN-inspired cellular-recurrent layers with bounded local processing, compact recurrent memory, routing, fusion, and accept-or-rollback validation. Three implementations are studied: Integrated Memory, MemoryFusion, and PDelta3-GDN2-CLVR+Local32. Strict PDelta3 conversion accepts a layer only when representation and NLL criteria pass fixed thresholds. On SmolLM2-135M, layers 0-2 are accepted with cumulative $Δ\mathrm{NLL}=+0.01209$, while layer 3 is rejected despite acceptable NLL because representation fidelity fails. On Qwen3.5-0.8B, full-attention layers 3, 7, and 11 are accepted with final $Δ\mathrm{NLL}=+0.02073$. Integrated Memory keeps perplexity within $-0.07\%$ to $+0.93\%$ while reducing total cache by up to $6.01\%$. A sampled 200-item downstream sanity check gives $28.5\%$--$32.0\%$ overall accuracy for converted Qwen releases. The results support conservative, quality-gated structural conversion rather than universal attention replacement or speedup.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑