arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越参数单体:语言模型的重构记忆、可执行技能与残差组装

Beyond the Parameter Monolith: Reconstructive Memories, Executable Skills, and Residual Assembly for Language Models

A. Bochkov

arXiv 2610.04012首次发表:更新:

AI 中文总结

本文提出FEM-ASM架构,将语言模型的存储、执行与神经协调分离,通过残差算子组装重构记忆和可执行技能,实验表明分离可行但重构精度和端到端效率仍有限。

AI 中文摘要

语言模型系统可以将上下文计算、持久存储和精确执行分离,而不是通过一个共享的参数系统更新所有能力。我们研究了FEM-ASM,一种受有限元方法启发的组织方式,其中独立构建的文档状态和确定性可执行技能向共享的语言模型状态贡献类型化提案。一个显式的残差算子协调附加到公共接口节点上的提案。我们通过受控实验和负面结果来评估这种组织方式,而不是声称语言具有物理上的有限元表述。一个无注意力的Multi-Mesh原型学习了因果语言建模,但并未建立有竞争力的通用能力。一个版本化存储包含52,809个重构记忆元素,接近17亿浮点值的预算;重构不完整,token准确率约为75%。支持感知的词汇索引使这些元素在受来源控制的查询构建下可寻址。对于可执行的算术,位置性结果观测相对于重复的全局结果向量显著改善了神经渲染,并且输出替换改变了模型的首选答案。一个有界附加演示进一步衡量了使选定证据可用的效果,而没有确定加载整个数十亿值存储的效用。结果支持存储、执行和神经协调的分离,同时指出了仅问题检索、无限制答案生成和端到端效率方面的未解决局限性。

英文摘要

Language-model systems can separate contextual computation, persistent storage, and exact execution instead of updating all capabilities through one shared parameter system. We investigate FEM-ASM, a finite-element-method-inspired organization in which independently constructed document states and deterministic executable skills contribute typed proposals to a shared language-model state. An explicit residual operator reconciles proposals attached to common interface nodes. We evaluate this organization through controlled experiments and negative results rather than claiming a physical finite-element formulation of language. An attention-free Multi-Mesh prototype learns causal language modeling but does not establish competitive general capability. A versioned store contains 52,809 reconstructive memory elements near a 1.7-billion-floating-value budget; reconstruction is incomplete, with approximately 75\% token accuracy. Support-aware lexical indices make these elements addressable under provenance-controlled query construction. For executable arithmetic, positional result observations substantially improve neural rendering relative to a repeated global result vector, and output substitutions change the model's preferred answer. A bounded attachment demonstration further measures the effect of making selected evidence available, without establishing the utility of loading an entire multi-billion-value store. The results support a separation of storage, execution, and neural coordination, while identifying unresolved limitations in question-only retrieval, unrestricted answer generation, and end-to-end efficiency.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑