发表机构
University of Oxford; Max Planck Institute for Software Systems; University of Birmingham(牛津大学; 马克斯·普朗克软件系统研究所; 伯明翰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出框架计算大语言模型生成有害输出概率的严格边界,利用克洛普 - 皮尔逊置信区间,通过潜在空间特征优先探索有害分支,能高效计算可靠下界,实验验证方法有效性,为模型评估和认证提供新途径。
AI 中文摘要
我们提出了一个新颖的框架,用于计算大语言模型(LLM)针对给定提示生成有害输出的概率的严格边界。我们研究了克洛普 - 皮尔逊置信区间的新应用,以获得该问题的概率近似正确(PAC)边界。作为主要技术贡献,我们提出一种算法,利用潜在空间中的特征,优先探索自回归生成树中更可能产生有害输出的分支。我们的方法尤其能高效计算有用的下界,即便真实危害概率极小,且所获下界是可靠的,即经形式证明小于实际危害概率。实验结果通过计算现有先进LLMs的非平凡下界证明了方法的有效性。本研究为LLMs的评估和统计认证提供了新途径。
英文摘要
We propose a novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to a given prompt. We study a new application of the Clopper-Pearson confidence intervals to obtain probably approximately correct (PAC) bounds for this problem. As our main technical contribution, we propose an algorithm that leverages features in the latent space to prioritize exploring branches in the auto-regressive generation tree that are more likely to produce harmful outputs. Our approach in particular enables the efficient computation of useful lower bounds, even in scenarios where the true harm probability is extremely small, and crucially, the obtained lower bounds are sound, i.e., formally proven to be less than the actual harmfulness probability: our experimental results demonstrate the effectiveness of our method by computing non-trivial lower bounds on state-of-the-art LLMs. This study newly enables the evaluation and statistical certification of LLMs.
CommentsThe Initial version of this manuscript has been available on OpenReview, see https://openreview.net/forum?id=papImkPLf5