迈向大语言模型安全知识的自动化测试
Semi-Automated Detection of Gaps in LLM Security Knowledge
- Northeastern University(东北大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究旨在实现大语言模型安全知识自动化测试,引入利用消费者保护机构权威信息识别LLM响应不稳定性的部分自动化方法,通过对身份盗窃和冒名顶替诈骗主题及Gemini和GPT家族中5个LLMs进行测试,可区分模型安全知识充足与否。
AI中文摘要:
大语言模型(LLMs)越来越多地用于一系列软件、硬件和以人类为中心的安全任务。因此,LLMs在安全任务上的性能是一个活跃的测量和研究领域,通常侧重于识别LLM安全“知识”可能不足的领域。识别LLM安全知识差距的常用策略包括构建挑战性问题语料库或任务基准,这些策略需要大量人工和安全专业知识来设计和执行。我们引入了一种部分自动化的方法来评估LLM对安全知识的掌握。该方法利用消费者保护机构(CPAs)的权威信息来识别LLM响应中的不稳定性,这些不稳定性可能表明存在知识差距。我们使用来自6个相关网站的关于身份盗窃和冒名顶替诈骗的公开信息,对2个安全主题(身份盗窃和冒名顶替诈骗)以及2个领先LLM家族(Gemini和GPT)中的5个LLMs演示了该方法。该方法能够区分具有足够知识和没有足够知识来准确识别文本叙述中安全主题的模型。
英文摘要:
Large language models (LLMs) are increasingly used for a range of software, hardware and human-centered security tasks. Consequently, LLM performance on security tasks is an active area of measurement and research, often with a focus on identifying areas in which LLM security "knowledge" may be insufficient. Popular strategies for identifying LLM security knowledge gaps include building corpora of challenge questions or task benchmarks, strategies that require substantial manual work and security expertise to design and execute. We introduce a partially-automated method for assessing LLM knowledge of a security area. The method uses authoritative information from Consumer Protection Agencies (CPAs) to identify instability in LLM responses that can be indicative of knowledge gaps. We demonstrate the method for 2 security topics, identity theft and impostor scams, and 5 LLMs in 2 leading LLM families, Gemini and GPT, using publicly available information about identity theft and impostor scams from 6 CPAs. The method distinguishes between models that have and don't have sufficient knowledge to accurately identify the security topics in text narratives.