arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言模型的安全纠错:冻结基座调整与能力保持

Safe Error Correction for Language Models: Frozen-Base Adjustment with Capability Preservation

Gautam Kishore

arXiv 2609.16145首次发表:更新:

发表机构

Eulogik(Eulogik)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出CRN v2,一个轻量级logit纠错模块,在冻结Gemma模型上通过SFT和DPO训练,以53.3%纠错率保持基础能力,验证了冻结基座加logit纠错加KL锚定的设计原则。

AI 中文摘要

我们研究一个实际问题:一个小型纠错模块能否修复冻结语言模型输出中的错误,同时不降低其基础能力?我们提出CRN v2,一个轻量级的logit级纠错模块(约3400万可训练参数,占46.5亿文本模块的0.73%),置于完全冻结的Gemma 4 E2B模型之上。基础模型从不更新;只有纠错模块通过监督微调及随后的无参考DPO在83,400个纠错对上进行学习。在60题领域考试(CEHRI:认证人机智能,涵盖事实、算术和隐式目标推理)中,CRN v2纠正了基础模型错误的53.3%(改写变体:43.3%),同时在测试的能力基准(MMLU/BoolQ N=200;洗车 N=8)上未显示退化。匹配CRN v1预算的LoRA基线(660万参数,秩19)实现了83.3%的纠错率,但在相同基准上遭受30-75%的能力损失——即纠错-能力权衡。消融实验表明,KL保持项(lambda=0.1)至关重要:将其降至0.01会使纠错率降至35.0%。在较早层进行隐藏状态注入的变体(160万参数,仅SFT)达到50.0%/55.8%,但未超过logit纠错;较浅注入(第4层)降至30.0%/28.3%;多深度logit纠错(约3500万)仅达到40%;更长的训练(5,000 SFT + 2,000 DPO)保持在53.3%——我们测试的替代配置均未超过秩128的logit结果,与约53%的最佳实现结果一致,而非下限。这是对设计原则(冻结基座+logit纠错+KL锚定)的研究,而非架构新颖性的声明。所有代码、主要结果权重和评估脚本均已发布(深度变体仅提供代码——无训练好的深度检查点)。

英文摘要

We study a practical question: can a small correction module fix errors in a frozen language model's outputs without degrading its base capabilities? We propose CRN v2, a lightweight logit-level correction module (~34M trainable parameters, 0.73% of the 4.65B text module) that sits atop a fully frozen Gemma 4 E2B model. The base model is never updated; only the correction module learns, via supervised fine-tuning followed by reference-free DPO on 83,400 error-correction pairs. On a 60-question domain exam (CEHRI: Certified Human-Robot Intelligence, covering facts, arithmetic, and implicit-goal reasoning), CRN v2 corrects 53.3% of base-model errors (reworded variant: 43.3%) while showing no degradation on tested capability benchmarks (MMLU/BoolQ N=200; car-wash N=8). A LoRA baseline at the matched CRN v1 budget (6.6M params, rank 19) achieves 83.3% correction but suffers 30-75% capability loss on the same benchmarks -- the correction-capability tradeoff. An ablation shows that the KL preservation term (lambda=0.1) is critical: lowering it to 0.01 degrades correction to 35.0%. A hidden-state injection variant at earlier layers (1.6M params, SFT-only) reaches 50.0%/55.8% but does not exceed logit correction; shallower injection (layer 4) drops to 30.0%/28.3%; multi-depth logit correction (~35M) reaches only 40%; and longer training (5,000 SFT + 2,000 DPO) stays at 53.3% -- none of the alternative configurations we tested exceeded the rank-128 logit result, consistent with a best-achieved result of ~53% rather than a floor. This is a study of a design principle (frozen base + logit correction + KL anchoring), not a claim of architectural novelty. All code, main-result weights, and evaluation scripts are released (deep variant as code only -- no trained deep checkpoints).

Comments10 pages, 4 tables. Code, weights, and evaluation scripts: https://github.com/eulogik/prajna and https://huggingface.co/eulogik/Prajna-CRNv2

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑