arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一个冻结的12B模型在经过验证的任务上超越前沿模型:100%准确率、零token、位精确、永远如此

A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever

Sietse Schelpe

arXiv 2607.23806首次发表:更新:

AI 中文总结

研究提出一种让模型冻结,通过增长持久内存来解决问题的方法。在多个问题族实例上,四种架构得分优异,执行与参数扩展解耦。该方法在开放式推理中也有效,内存选择快,存储规模大,在已验证任务上超越前沿模型。

AI 中文摘要

如今改进语言模型意味着重新训练:需要巨大计算量,每个周期都有新的不透明模型,输出不确定。我们走相反路径:模型保持冻结,在其旁边增长经过验证的解决方案的持久内存。一旦一个问题族得到解决并通过独立验证步骤(从不参考答案密钥),该问题族的每个新实例都能在零生成token时被确定地、位精确地回答。在跨越九个问题族的180个新实例上,来自四个供应商的四种架构(密集型和专家混合型)在每个答案的零生成token时得分均为180/180:执行能力与参数扩展解耦。一个负控制将能力完全归因于内存:清空内存后,它什么也解决不了。对于开放式推理,相同的先验证后存储契约也适用:所有四个模型的一致性门控接受率为88/88,机器检查形式证明,推理方法转移率为77/80。内存选择耗时1.4微秒;一次完整重用在6 - 23毫秒内完成,功耗为36毫瓦。在一个有4500个项目的经过验证的存储中,近似相似性检索选错项目的概率为94.3%,而精确寻址零错误。该存储还作为工作上下文,规模是任何已发布引擎都无法比拟的:在单个46GB GPU上有一个6000000token的可移动窗口,而vLLM在30399token时停止,SGLang在超过32000token时会静默截断。在已发布的基准测试中,前沿模型在从头开始的原始推理方面仍远超任何12B模型;但在这个系统已解决和验证的所有方面,情况相反:前沿API调用每次查询都要进行新的生成过程,而经过验证的重用每次都零token成本并返回相同的位。本报告附带一个免费的、有速率限制访问的公共测试平台:此https URL

英文摘要

Improving a language model today means retraining it: enormous compute, a new opaque model each cycle, non-deterministic output. We take the opposite path: the model stays frozen, and a persistent memory of verified solutions grows beside it. Once a problem family is solved and has passed an independent verification step that never consults the answer key, every new instance of that family is answered at zero generation tokens, bit-exact, deterministically. Across 180 fresh instances spanning nine problem families, four architectures from four vendors - dense and mixture-of-experts - each score 180/180 at zero generation tokens per answer: execution-bound capability decoupled from parameter scaling. A negative control attributes the capability fully to the memory: emptied, it solves nothing. The same verify-before-store contract holds for open-ended reasoning: 88/88 consistency-gated acceptances across all four models, machine-checked formal proof, and reasoning-method transfer at 77/80. Memory selection takes 1.4 microseconds; a full reuse completes in 6-23 ms at 36 mWh. Approximate similarity retrieval selects the wrong item 94.3% of the time on a 4,500-item verified store where exact addressing makes zero errors. The store also serves as working context at a scale no shipped engine matches: a 6,000,000-token movable window on a single 46 GB GPU at flat memory, where vLLM stops at 30,399 tokens and SGLang silently truncates past 32,000. On published benchmarks, frontier models remain far ahead of any 12B at raw from-scratch reasoning; on everything this system has solved and verified, the comparison inverts: a frontier API call pays a fresh generation pass on every query, forever, while verified reuse costs zero tokens and returns the identical bits every time. A public testbench with free, rate-limited access accompanies this report: https://corbenic-galahad-bench.hf.space

CommentsIndustry experience report. 14 pages, 8 figures. Public testbench: https://corbenic-galahad-bench.hf.space; companion repository with SHA-256 provenance manifest: https://github.com/corbenicai/galahad-bench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑