arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FORGE:生成式语言模型的验证门控行为修复

FORGE: Verification-Gated Behavioral Repair for Generative Language Models

Hsin-Ling Hsu, Min-Yu Chen, Nai-Chia Chen, Yan-Ru Chen, Yi-Ling Chang, Fang Yu

arXiv 2610.05190首次发表:更新:

发表机构

National Chengchi University(国立政治大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FORGE提出一种验证门控的生成式语言模型行为修复框架,通过分离缺陷定位、权重编辑和行为验证,结合约束优化与零空间投影两种后端,在少量缺陷样本下有效降低偏见和毒性,优于梯度微调。

AI 中文摘要

生成式大语言模型(LLMs)从预训练中继承了不良行为,包括人口统计偏见和有毒内容生成,这些行为通常仅在部署后出现,并影响一小部分输入。修复应消除已识别的缺陷,保持模型的整体功能,并理想情况下提供正确性保证。现有方法仅部分解决了这一问题:基于梯度的微调缺乏逐实例保证,且在缺陷样本较少时变得不稳定;模型编辑假设显式知识替换而非行为纠正;基于约束的修复主要局限于具有唯一目标输出的判别式模型。我们提出了FORGE,一个针对生成式语言模型的定向行为修复框架,将缺陷定位、权重编辑和行为验证分离为独立阶段。其核心是一种修复抽象,将局部缺陷生成转化为显式优化目标,使面向验证的修复技术能够应用于自回归生成。FORGE与编辑机制无关:我们通过(1)一种基于约束的二次优化方法(提供逐样本修复证书)和(2)一种零空间投影编辑器(最小化对原始模型分布的干扰)来实例化它,两者均在相同的定位和验证协议下运行。在五个开源LLM上,FORGE在减少偏见和毒性方面始终比基于梯度的微调取得更大降幅,且困惑度退化较小。两种后端在不同架构上表现出互补性能,轻量级因果探针将其追溯到毒性相关信号的集中位置。FORGE在仅有少量缺陷示例时也保持有效,而传统微调常常振荡或无法收敛。

英文摘要

Generative large language models (LLMs) inherit undesirable behaviors from pre-training, including demographic bias and toxic generation, that often emerge only after deployment and affect a small subset of inputs. A repair should eliminate the identified defect, preserve the model's overall functionality and, ideally, provide correctness guarantees. Existing approaches address this only partially: gradient-based fine-tuning lacks per-instance guarantees and becomes unstable with few defect samples; model editing assumes explicit knowledge replacement rather than behavioral correction; and constraint-based repair is largely restricted to discriminative models with unique target outputs. We present FORGE, a framework for targeted behavioral repair of generative language models that separates defect localization, weight editing, and behavioral verification into independent stages. Its core is a repair abstraction that converts localized defective generation into explicit optimization objectives, enabling verification-oriented repair techniques to operate on autoregressive generation. FORGE is editing-mechanism agnostic: we instantiate it with (1) a constraint-based quadratic optimization method that provides per-sample repair certificates and (2) a null-space projection editor that minimizes interference with the original model distribution, both under the same localization and verification protocol. On five open-source LLMs, FORGE consistently achieves larger reductions in bias and toxicity than gradient-based fine-tuning with minor perplexity degradation. The two backends exhibit complementary performance across architectures, which a lightweight causal probe traces to where toxicity-related signals concentrate. FORGE also remains effective with only a handful of defective examples, where conventional fine-tuning often oscillates or fails to converge.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑