一种用于大语言模型护栏的双假设推理框架
A Dual-Hypothesis Reasoning Framework for LLM Guardrails
浏览论文内容
中文总结 AI 辅助
该研究提出ARBITER大语言模型护栏框架,采用双假设推理和多组件监督微调方法,用成本效益高的策略生成推理痕迹及基于LoRA微调,性能超现有方法,还能为不安全决策提供解释,在安全审核基准实验中表现出色。
中文摘要 AI 辅助
我们提出了ARBITER,一种新颖的大语言模型护栏框架,它引入了两个关键思想:(i)双假设推理,一种用于大语言模型护栏的推理方法,在做出安全决策之前明确考虑提示的安全和不安全解释;(ii)多组件监督微调(MC-SFT),一种基于推理的护栏的结构化训练损失,将大语言模型输出分解为逻辑组件,并根据其重要性加权。现有基于推理的护栏通常依赖于昂贵的程序。相比之下,ARBITER使用具有成本效益的自我生成策略来生成推理痕迹,并基于LoRA进行参数高效微调,同时性能优于这些昂贵方法。此外,ARBITER为不安全决策提供可靠的证据短语解释。在三个安全审核基准上的实验表明,ARBITER优于现有的基于推理和非推理的护栏基线,在域外评估中有明显提升。
英文摘要
We propose ARBITER, a novel LLM guardrail framework that introduces two key ideas: (i) dual-hypothesis reasoning, a reasoning method for LLM guardrails that explicitly considers both safe and unsafe interpretations of a prompt before making a safety decision, and (ii) multi-component supervised fine-tuning (MC-SFT), a structured training loss for reasoning-based guardrails that decomposes LLM outputs into logical components and weights them according to their importance. Existing reasoning-based guardrails often rely on expensive procedures, such as generating reasoning traces using larger or closed-source teacher models and applying full-parameter fine-tuning. In contrast, ARBITER uses a cost-effective self-generation strategy for reasoning traces and LoRA-based parameter-efficient fine-tuning while still achieving better performance than these expensive approaches. Additionally, ARBITER provides faithful evidence-phrase explanations for unsafe decisions, enabling a more transparent and interpretable guardrail method. Experiments on three safety moderation benchmarks show that ARBITER outperforms existing reasoning-based and non-reasoning guardrail baselines, with clear gains in out-of-domain evaluations.
发表机构
- University of Arizona(亚利桑那大学)
机构由 AI 辅助整理,请以论文原文为准。