Auto-Tuning Safety Guardrails for Black-Box Large Language Models
为黑盒大语言模型自动调整安全防护栏
机构 * University of St. Thomas(圣托马斯大学)
专题命中 越狱攻击 :safety(title,abstract);jailbreak(abstract);分类 cs.CL、cs.LG
AI总结 本文提出通过超参数优化来自动调整黑盒大语言模型的安全防护栏,以提高其安全性和效率。
Comments 8 pages, 7 figures, 1 table. Work completed as part of the M.S. in Artificial Intelligence at the University of St. Thomas using publicly available models and datasets; all views and any errors are the author's own