AI 中文总结
本文提出LatentPort方法,首次实现不同规模混合语言模型间无需重放上下文的持久循环状态迁移,通过翻译KV加GDN状态包显著降低NLL,并用小参数修正逼近原生模型性能。
AI 中文摘要
一个语言模型能否将其实时记忆移交给另一个模型,而接收方无需重新阅读上下文?我们演示了在一对架构匹配的Qwen3.5 4B到9B兄弟模型之间进行有用的持久混合状态迁移。据我们所知,这是首次在不同规模的混合语言模型之间演示无需目标前缀重放的持久循环推理状态的跨模型交接。仅翻译注意力KV会留下很大差距;添加Gated DeltaNet(GDN)持久状态包将教师强制的负对数似然(NLL),即平均下一个词元的对数损失,降低了0.747 nats/词元(95%配对文档自助法置信区间[0.6921, 0.8047]),改善了全部64个PG19文档。直接循环和卷积重用优于测试的学习型GDN映射,这与持久状态坐标的部分功能兼容性一致。一个全新的组件因子分析选择了翻译的KV与直接循环和卷积状态。一个额外的434,176参数修正改进了该基础方法在64个全新网络文档上的表现:延续损失比原生9B高0.076 nats/词元(超额NLL),Jensen-Shannon(JS)散度为0.022,原生上下文恢复(NCR)为0.918。修正后的9B在处理零历史前缀词元的情况下显著优于继续4B推理。证据覆盖一个方向、一对几何匹配的基础模型对以及4K教师强制延续;接近原生的门控失败,16K分支未运行,自由生成等价性和通用状态接口仍未得到证实。
英文摘要
Can one language model hand its live memory to another without the receiver rereading the context? We demonstrate useful persistent hybrid-state transfer across one architecture-matched Qwen3.5 4B-to-9B sibling pair. To our knowledge, this is the first demonstrated cross-model handoff of persistent recurrent inference state between differently sized hybrid language models without target prefix replay. Translated attention KV alone leaves a large gap; adding the Gated DeltaNet (GDN) persistent-state package lowers teacher-forced negative log-likelihood (NLL), the average next-token log-loss, by 0.747 nats/token (95% paired document bootstrap CI [0.6921, 0.8047]), improving all 64 PG19 documents. Direct recurrent and convolution reuse outperforms the tested learned GDN maps, consistent with partial functional compatibility of persistent-state coordinates. A fresh component factorial selects translated KV with direct recurrent and convolution state. An additional 434,176-parameter correction improves that base on 64 fresh web documents: continuation loss is 0.076 nats/token above native 9B (excess NLL), Jensen-Shannon (JS) divergence is 0.022, and native context recovery (NCR) is 0.918. Corrected 9B significantly beats continued 4B inference while processing zero historical prefix tokens. Evidence covers one direction, one geometry-matched Base-model pair, and 4K teacher-forced continuation; the near-native gate failed, the 16K branch was not run, and free-generation equivalence and a general state interface remain unproven.
Comments14 pages, 6 figures, 12 tables