发表机构
Institute for Frontier Interdisciplinary Research in Health Sciences and Technology, Sun Yat-sen University; Guangdong Engineering Research Center of Medical Artificial Intelligence Multimodal System; Sun Yat-sen University; School of Business, Sun Yat-sen University; School of Computer Science, China University of Geosciences (Wuhan); School of Public Health, Sun Yat-sen University; Hospital of Stomatology, Sun Yat-sen University; Big Data and Artificial Intelligence Center, The Third Affiliated Hospital of Sun Yat-sen University(中山大学健康科学与技术前沿交叉研究院; 广东省医学人工智能多模态系统工程研究中心; 中山大学; 中山大学商学院; 中国地质大学(武汉)计算机学院; 中山大学公共卫生学院; 中山大学口腔医院; 中山大学附属第三医院大数据与人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CellFill方法通过在量化决策单元内写入残差实现量化大语言模型的比特一致、可撤销更新,实验显示其在召回率接近基准的同时降低跨域遗忘,且可迁移至27B混合线性注意力模型。
AI 中文摘要
对已部署语言模型进行新知识更新的所有现有方式(全微调、适配器合并、模型编辑)都会替换已发布的检查点,进而替换所有引用这些精确比特的评估和缓存。本文提出的CellFill方法则在反量化间隙中学习:冻结4位发布模型的整数编码和缩放因子,仅将新知识写入严格位于每个量化决策单元内的每个权重残差中。随后重新量化可返回与发布工件完全一致的比特,这是可通过机器验证的保证;更新可通过丢弃残差实现完全撤销;且漂移是受限的。本文给出6个命题和3种训练路径,包括CellFill——一种受限重参数化方法,其将不变性构建为结构而非强制约束。实验结果显示,精确不变性几乎是无代价的:在3组配对随机种子下,受限密集路径的事实召回率为58.9%,与权重可脱离工件的无约束基准(59.3%)接近,配对差值为-0.5个百分点,95%置信区间为[-5.0,+4.0];且在保留的跨域困惑度上表现更优。针对零假设(将相同更新作为未合并适配器提供服务),投影到单元内在所有收敛的运行中均降低了跨域遗忘,而发散的对照实验则显示了该方法的边界:投影是一个信赖域,而非修复。没有任何方法能规避知识本身的成本,而在域内困惑度上超过锚点的看似免费午餐是排练与指标共享语料库的产物。在匹配的排练设置下,不同方法在每单位跨域困惑度所获取的知识上存在3倍差异,该排名与召回率排名不同。该方法可迁移到27B混合线性注意力模型(2.4e10个受限权重,已验证比特一致),在该模型上,匹配召回率的跨域困惑度成本约为1.7B模型的一半。
英文摘要
A 4-bit quantized weight specifies a rounding cell rather than a single full-precision value. We introduce in-cell learning, a paradigm for writing new knowledge only within these cells, so that re-quantizing the served weights reproduces the released integer codes and scales exactly. CellFill implements this idea with bounded trainable positions inside frozen quantization cells and ships the update as a separate, subtractively revocable file. Across published NF4 and W4A16 releases of Qwen3 and Gemma from 1.7B to 32B parameters, CellFill writes 83-99% of a real-fact corpus while returning the stored code on every constrained weight. The injected facts generalize to paraphrases and composition, and answer 78-88% of selected PopQA questions that the released model misses. Sequential experiments show that rehearsal preserves earlier knowledge, whereas available room and new-task plasticity decline across updates. Consolidation re-quantizes the learned weights to produce a declared major version, restoring room at a measured capability cost. A six-task write-rehearse-consolidate cycle retains at least 92.8% of first learning in two 8B runs and records zero code violations over 6.9 billion constrained weights at every fold. These results define a version-management protocol in which minor updates preserve the released quantized artifact bitwise and major updates are explicit, measurable, and verifiable.
Comments94 pages, 14 figures, 22 tables. Technical Report, version 3. Code and archived artifacts: https://github.com/sumsliu/in-cell-learning