arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33638cs.LGcs.CLstat.ML

量化黑盒语言模型中的行为尾部

Quantifying Behavioral Tails in Black-Box Language Models

Elsayed Eshra, Ali Al-Lawati, Dongwon Lee, Suhang Wang

首次发表
浏览论文内容

中文总结 AI 辅助

RareTrap框架通过几何感知映射和顺序稀有事件模拟,仅用200次评估即可估计黑盒LLM中严重行为的概率,为模型安全评估提供原则性方法。

中文摘要 AI 辅助

我们提出了RareTrap,一个用于估计黑盒大语言模型(LLM)中严重行为概率的框架。概率估计的一个关键挑战是定义输入空间上的可处理分布。为实现这一目标,RareTrap使用一个替代LLM,并构建从低维潜在参考空间到其词元嵌入空间的几何感知映射,以在输入提示上诱导出明确且可复现的分布。利用响应级性能函数对响应进行量化,以评估行为严重性。这使得顺序稀有事件模拟能够将评估集中在逐渐更严重的行为上,同时保持诱导提示分布下的概率,否则这将难以测量。在10个开放权重模型和两个前沿模型(GPT-5.4和Claude Sonnet 4.6)上,我们发现RareTrap成功诱导了严重的资源消耗行为,并仅用200次评估就计算了其概率。RareTrap为模型开发者提供了一种在共同分布下评估语言模型的原则性方法,并优先调整对齐工作以提升安全性和降低风险。代码已在线发布:此https URL。

英文摘要

We introduce RareTrap, a framework for estimating the probability of severe behaviors in black box large language models (LLMs). A key challenge for probability estimation is defining a tractable distribution over the input space. To accomplish that, RareTrap uses a surrogate LLM and constructs a geometry-aware mapping from a lower-dimensional latent reference space into its token-embedding space to induce an explicit and reproducible distribution over input prompts. A response-level performance function is utilized on the response to quantify behavior severity. This enables sequential rare event simulation that concentrates evaluations on progressively more severe behaviors while preserving probability under the induced prompt distribution, which would otherwise be prohibitive to measure. Across 10 open-weight and two frontier models (GPT-5.4 and Claude Sonnet 4.6), we find that RareTrap successfully induces severe resource consumption behaviors and computes their probability with as few as 200 evaluations. RareTrap provides model developers a principled approach for evaluating language models under a common distribution, and prioritizing alignment effort to improve safety and mitigate risks.

发表机构

  • The Pennsylvania State University(宾夕法尼亚州立大学)

机构由 AI 辅助整理,请以论文原文为准。

↑