发表机构
JPMorgan(摩根大通)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究使用闭源大语言模型的实际应用的安全措施问题,提出用在合成数据上训练的小型语言模型作专用护栏,引入受GAN启发的合成数据生成方法,实验证明该方法训练的护栏比基于提示的性能更优。
AI 中文摘要
使用闭源大语言模型(LLM)的实际应用需要超越基本内容过滤器的高级安全措施。内容审核过滤器(如毒性和偏差)有相对标准的定义,而诸如幻觉、主题漂移和行为偏差等特定于应用的护栏更难建模且因用例而异。此外,数据稀缺和标注成本使得创建和测试专用护栏具有挑战性。在这项工作中,我们提出使用在合成数据上训练的小型语言模型(SLM)作为LLM应用的专用护栏。我们引入了一种受生成对抗网络(GAN)设计启发的新型合成数据生成方法,以生成高质量的合成数据样本,可用于训练SLM来编码特定于用例的护栏信息,从而充当专用护栏。我们的实验表明,在高质量合成数据上训练的SLM护栏比基于提示的LLM护栏性能更优。
英文摘要
Real-world applications that use closed-source large language models (LLMs) need advanced safety measures that go beyond the basic content filters. Content moderation filters such as toxicity and bias have relatively standard definitions where as application specific guardrails like hallucination, topic drift and behaviour deviation are more difficult to model and can vary by use case. Additionally, data scarcity and annotation costs, make the process of creating and testing specialized guardrails challenging. In this work, we propose using Small Language Models (SLMs) trained on synthetic data as specialized guardrails for LLM applications. We introduce a novel synthetic data generation method inspired by the design of Generative Adversarial Networks (GANs) to generate high quality synthetic data samples which can be used to train SLMs to encode use case specific guardrail information and hence function as specialized guardrails. Our experiments demonstrate that SLM guardrails trained on high quality synthetic data show performance gains over prompt based LLM guardrails.