发表机构
Corbenic AI(Corbenic人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出通过字节精确的键值缓存嫁接,在不改变权重的情况下让冻结小语言模型更强大且廉价。此方法能精确恢复知识,提升模型性能,如在AIME 2025上提高准确率,还能扩展可用上下文,减少能耗,且可验证评分。
AI 中文摘要
我们报告了一种方法,能让冻结的小语言模型在不改变任何权重的情况下,同时变得更强大且成本大幅降低。经过验证的知识作为字节精确的键值(KV)状态工件一次性存入,随后通过嫁接恢复到新的推理上下文。恢复是位精确的:在固定的确定性配置下,嫁接后的逻辑输出与全新计算的字节逐字节相同(SHA - 256相等),在五十个样本上KL散度为零且argmax一致性达100%。我们表明自身位置嫁接是具有浮点旋转编码的模型上唯一数值精确的操作点,并在两个模型规模(12B、31B)和两个GPU目标上验证了字节精确性,其中一个通过预注册重放。在AIME 2025上,一旦嫁接经过验证的解决方案库,冻结的Gemma - 4 - 12B从80.0%提升到93.3%,高于其自身的77.5%以及其31B同级的89.2%已发布基准。在重复案例中,基础模型在401,026令牌预算内从未解决的八个问题,从缓存的经过验证的解决方案中在总共61个解码令牌中得到回答,令牌数减少了6574倍,能量消耗减少了约8700倍;能力声明基于留出的转移(31B时7个中有7个)。相同的字节精确存储在不增加额外加速器内存的情况下将可用上下文从32768扩展到2854766令牌,并在相同架构的机器之间实现字节相同的移动。我们在行为层面描述了该系统;引擎是专有的,每个报告的数字都由提交的输入和输出哈希支持,因此无需引擎也可重新检查评分。
英文摘要
We report a way to make a frozen small language model both more capable and dramatically cheaper at once, without changing any weights. Verified knowledge is deposited once as a byte-exact key-value (KV) state artifact and later restored, by graft, into a fresh inference context. The restore is bit-exact: under a pinned deterministic configuration, the grafted logits are byte-for-byte identical to a fresh computation (SHA-256 equality), with zero KL divergence and 100% argmax agreement over fifty samples. We show that own-position graft is the unique numerically exact operating point on a model with floating-point rotary encoding, and we verify byte-exactness on two model scales (12B, 31B) and two GPU targets, one through a pre-registered replay. On AIME 2025, a frozen Gemma-4-12B moves from 80.0% to 93.3% once a verified solution library is grafted, above its own 77.5% and its 31B sibling's 89.2% published anchors. On the recurring case, eight problems the base model never solves within a 401,026-token budget are answered from cached verified solutions in 61 total decode tokens, a factor of 6,574 fewer tokens and about 8,700x less energy; the capability claim proper rests on held-out transfer (7 of 7 at 31B). The same byte-exact store widens usable context from 32,768 to 2,854,766 tokens at zero extra accelerator memory, and moves byte-identical between machines of the same architecture. We describe the system at the behavior level; the engine is proprietary, and every reported number is backed by committed input and output hashes so the scoring can be re-checked without it.
Comments18 pages, 4 figures