DataShield:通过共识子空间对齐揭示跨语言模型的风险微调数据
DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment
浏览论文内容
中文总结 AI 辅助
研究针对LLMs微调数据存在安全风险及现有方法局限性,提出DataShield框架,通过共识子空间对齐识别风险微调样本和响应段,经实验对比,该方法能有效降低误识率,还保留下游效用并避免特定模型风险计算。
中文摘要 AI 辅助
在特定领域数据集上微调大语言模型(LLMs)已成为使LLMs适应特定应用的标准范式。但近期研究表明,即使在良性任务特定数据上微调也会大幅削弱LLMs的安全能力。现有方法在识别导致安全降级的数据方面取得了进展,但通常依赖单个平均向量来表示安全方向,限制了风险评估措施的有效性和可转移性。为解决这些限制,我们提出DataShield,一个通过在多个安全对齐的LLMs导出的联合安全关键语义空间上进行共识子空间对齐来识别风险微调样本和响应段的数据评估框架。在这些空间中,DataShield使用安全和不安全数据表示的语义谱分解来提取共识安全和不安全子空间。然后通过测量数据样本或段与不安全和安全子空间的相对对齐来估计其风险,实现样本级过滤和细粒度段级掩码。与现有过滤和掩码基线相比,DataShield在样本过滤时将误识率降低了14.6%,在段掩码时降低了32.3%,同时保留了下游效用并避免了特定目标模型的风险计算。
英文摘要
Fine-tuning large language models (LLMs) on domain-specific datasets has become a standard paradigm for adapting LLMs to specialized applications. However, recent work has shown that even fine-tuning on benign task-specific data can substantially weaken the safety capabilities of LLMs. While existing efforts have made progress in identifying data responsible for safety degradation, they usually rely on a single mean vector computed over a specific model with its tokenizer to represent the safety direction, which limits both the effectiveness and transferability of their risk assessment measures. To address these limitations, we propose DataShield, a data assessment framework that identifies risky fine-tuning samples and response segments through consensus subspace alignment over joint safety-critical semantic spaces derived from multiple safety-aligned LLMs. Within these spaces, DataShield extracts consensus safe and unsafe subspaces using semantic spectral decomposition over safe and unsafe data representations. The risk of a data sample or segment is then estimated by measuring its relative alignment with the unsafe and safe subspaces, enabling both sample-level filtering and fine-grained segment-level masking. Compared with state-of-the-art filtering and masking baselines, DataShield reduces ASR by 14.6\% with sample filtering and 32.3\% with segment masking, while preserving downstream utility and avoiding target-model-specific risk computation.