arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SafeTune:一个用于审计和修复微调大语言模型中安全漂移的统一忠实库

SafeTune: A Unified Faithful Library for Auditing and Repairing Safety Drift in Fine-Tuned LLMs

Pratinav Seth, Saisab Sadhu, Anshul Kaushal, Vinay Kumar Sankarapu

arXiv 2609.22153首次发表:更新:

发表机构

Lexsi Labs(Lexsi 实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SafeTune是一个统一库,通过四种干预范式审计和修复微调LLM的安全漂移,提供配置驱动工作流和模块化注册表,并在金融和医疗案例中验证其有效性。

AI 中文摘要

针对微调大语言模型(LLMs)中安全漂移的解决方法分散在不兼容的实现、生命周期阶段和评估协议中,使得它们难以采用和比较。我们引入了SafeTune,一个源代码可用的库,它统一了四种干预范式:事后权重恢复、安全约束微调、基于梯度的遗忘以及推理时引导,并共享可解释性、评估和部署工具。SafeTune提供了一致的配置驱动工作流,同时保留了每种范式所需的独特输入和干预点。其模块化注册表支持新方法、基准、评判器、模型和微调领域,而无需重新设计周围的流程。我们通过受控比较以及金融和医疗部署案例研究展示了SafeTune,说明了它如何表征安全漂移,评估常见拒绝行为和能力评估上的可行干预措施,并支持校准或分层缓解。

英文摘要

Methods for addressing safety drift in fine-tuned Large Language Models (LLMs) are scattered across incompatible implementations, lifecycle stages, and evaluation protocols, making them difficult to adopt and compare. We introduce SafeTune, a source-available library that unifies four intervention paradigms: post-hoc weight recovery, safety-constrained fine-tuning, gradient-based unlearning, and inference-time steering, alongside shared interpretability, evaluation, and deployment utilities. SafeTune provides a consistent configuration-driven workflow while preserving the distinct inputs and intervention points each paradigm requires. Its modular registry supports new methods, benchmarks, judges, models, and fine-tuning domains without redesigning the surrounding pipeline. We demonstrate SafeTune through controlled comparisons and finance and medical deployment case studies, showing how it characterizes safety drift, evaluates feasible interventions on common refusal-behavior and capability evaluations, and supports calibrated or layered mitigation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑