arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14644cs.LGcs.CL

DUET:基于同权重分歧的双教师在线策略蒸馏用于禁止合规性

DUET: Dual-Teacher On-Policy Distillation via Same-Weight Disagreement for Prohibition Compliance

Zihan Li, Feifei Li, Wenhui Que

AI总结:

本研究针对LLM部署中的动态禁止规则合规问题,提出DUET双教师在线策略蒸馏方法,构建工业基准,在Qwen模型上实现高合规性与效用保留,性能优于基线。

AI中文摘要:

现实世界中的大语言模型(LLM)部署越来越依赖于运行时注入的禁止规则——企业政策、个人身份信息(PII)红线、工具边界等,这些规则会随每个请求和每个租户而变化。传统的后训练方法在结构上不适用:监督微调(SFT)将违规信号隐藏在合规标签中,而直接偏好优化(DPO)的序列级偏好与令牌局部违规不匹配。我们提出DUET,一种用于禁止合规性的令牌选择性在线策略蒸馏方法。DUET将一个可见禁止规则的教师(正教师)与一个权重完全相同但不可见禁止规则的教师(负教师)配对。由于两个教师仅在禁止规则的可见性上存在差异,它们的逐令牌分歧可分离出禁止规则的因果效应——产生不受模型容量或不匹配污染的干净监督信号。这种分歧驱动两种互补机制:信号清理,将一致令牌视为冗余或前缀损坏而丢弃;偏好导向学习,在令牌粒度上将学生模型推离负教师并推向正教师,将DPO风格的优化直接嵌入在线策略蒸馏(OPD)中,无需离线偏好数据。我们构建了一个工业级禁止合规性基准,涵盖五个任务系列,包括明确拒绝、释义鲁棒性和过度拒绝。在15亿至80亿参数的通义千问(Qwen)变体模型上,DUET实现了72.3%至85.2%的违规合规性,同时保留了88%至93%的正常效用,显著优于教师模型和其他蒸馏基线方法。对SysBench的外部评估证实,其安全对齐性得到提升,且在GSM8K和MATH-500上的性能下降极小。

英文摘要:

Real-world LLM deployments increasingly rely on runtime-injected prohibitions--enterprise policies, PII redlines, tool boundaries--that vary per request and per tenant. Conventional post-training is structurally ill-suited: SFT hides the violation signal in compliant labels, and DPO's sequence-level preferences mismatch token-localized violations. We propose DUET, a token-selective on-policy distillation method for prohibition compliance. DUET pairs a teacher that sees the prohibition (positive) with an identical-weight teacher that does not (negative). Because the two teachers differ only in prohibition visibility, their per-token disagreement isolates the prohibition's causal effect--yielding a clean supervision signal uncontaminated by model capacity or mismatch. This disagreement drives two complementary mechanisms: signal cleaning, which discards agreement tokens as redundant or prefix-corrupted, and preference-directed learning, which pushes the student away from the negative teacher and toward the positive one at token granularity, embedding DPO-style optimization directly into OPD without offline preference data. We construct an industrial Prohibition-Compliance benchmark spanning five task families covering explicit-refusal, paraphrase robustness, and over-refusal. Across 1.5B-8B Qwen variants, DUET achieves 72.3-85.2% violation compliance while preserving 88-93% normal utility, dramatically outperforming teacher model and other distillation baselines. External evaluation on SysBench confirms improved safety alignment with minimal degradation on GSM8K and MATH-500.

↑