AI 中文总结
本研究提出基于GFlowNets的自动化自适应红队评估方法,训练攻击者模型测试目标大语言模型,可生成更有效的英文攻击及土耳其语攻击输入,用于评估模型鲁棒性。
AI 中文摘要
大语言模型(LLMs)的快速发展推动其在各领域广泛应用,然而这一趋势也带来了严重的安全漏洞,亟需识别并缓解恶意利用产生的缺陷。红队评估通过多样的对抗输入评估模型鲁棒性,对暴露安全风险并制定应对措施至关重要。当前红队评估要么由专家手动执行,要么使用预定义攻击数据集自动执行,但手动测试耗时,现有自动方法因依赖固定数据集而创造力有限。本研究提出一种自动化、无需人工干预的自适应方法,利用GFlowNets,通过一个大语言模型测试另一个以识别LLM漏洞。在该框架中,攻击者模型针对指定目标模型进行训练,以执行自动化红队评估并提供量化鲁棒性分数。本研究旨在生成比现有基准更有效的英文对抗攻击,且作为文献的新贡献,引入了一种能生成土耳其语攻击输入的模型。
英文摘要
The rapid advancement of Large Language Models (LLMs) has facilitated their ubiquitous integration into various domains, leading to widespread adoption. However, this escalating trend has introduced significant security vulnerabilities, necessitating the identification and mitigation of flaws arising from malicious exploitation. Red teaming assessments, conducted to evaluate model robustness through diverse adversarial inputs, are essential for exposing security risks and implementing countermeasures. Currently, red teaming is performed either manually by experts or automatically using predefined attack datasets. Nevertheless, manual testing remains time-consuming, while existing automated methods suffer from limited creativity due to their inherent dependency on fixed datasets. In this study, we propose an automated, human-independent, and adaptive approach leveraging GFlowNets to identify LLM vulnerabilities by utilizing one large language model to test another. Within this framework, an attacker model is trained against a specified victim model to perform automated red teaming and provide a quantitative robustness score. This research aims to generate more effective adversarial attacks in English compared to existing benchmarks and, as a novel contribution to the literature, introduces a model capable of generating attack inputs in the Turkish language.