用适当评估估计语言模型中的罕见事件
Estimating Rare Events in Language Models with Proper Evaluation
浏览论文内容
中文总结 AI 辅助
研究语言模型中罕见事件概率估计难题,提出梯度激活自适应多级分裂方法及移位幂布雷格曼损失,通过实验揭示偏差 - 方差权衡,强调估计器与部署上下文匹配,确立激活空间为易处理域。
中文摘要 AI 辅助
量化语言模型中罕见故障的风险,如对抗性分布变化或大规模部署引发的故障,需要估计对随机采样而言过小的概率。近期工作虽已将低概率估计形式化,但现有管道在最罕见情况下仍很脆弱。本文引入梯度激活自适应多级分裂(GA - AMLS),将罕见事件蒙特卡罗方法应用于语言模型的连续激活空间。具体而言,GA - AMLS使用基于梯度的MCMC内核导航激活空间,消除输入空间搜索的零估计崩溃,并以显式、重尾激活先验下的条件采样取代先前激活空间估计器的独立性假设。还提出移位幂布雷格曼(SPB)损失,这是一种适当的评分规则,对零估计保持有限,并在低估和高估惩罚之间提供可调不对称性。在小型变压器模型上的实验揭示了偏差 - 方差权衡:GA - AMLS在对称评估下实现最低损失,相对于所有模型大小的最强基线降低平均对数空间平方误差,而在不对称惩罚下具有高估偏差的方法占优。研究结果强调估计器选择应与部署上下文匹配。更广泛地说,本文工作将激活空间确立为语言模型中罕见事件估计的易处理域,规避离散输入空间搜索的脆弱性。
英文摘要
Quantifying the risk of rare failures in language models, such as those triggered by adversarial distribution shifts or very large-scale deployments, requires estimating probabilities far too small for random sampling. While recent work has formalized Low Probability Estimation, existing pipelines remain fragile in the rarest regimes: estimators can suffer zero-estimate collapse or systematic bias, and standard evaluation losses can become unstable or poorly matched to asymmetric safety costs. In this work, we introduce Gradient Activation Adaptive Multi-Level Splitting (GA-AMLS), which adapts rare-event Monte Carlo methods to the continuous activation space of language models. Specifically, GA-AMLS uses a gradient-based MCMC kernel to navigate activation space, eliminating the zero-estimate collapse of input-space search and replacing the independence assumptions of prior activation-space estimators with conditional sampling under an explicit, heavier-tailed activation prior. We also propose the Shifted-Power Bregman (SPB) Loss, a proper scoring rule that remains finite for zero-estimates and offers tunable asymmetry between underestimation and overestimation penalties. Experiments on small transformer models reveal a bias-variance tradeoff: GA-AMLS achieves the lowest loss under symmetric evaluation, reducing average log-space squared error relative to the strongest baseline across model sizes, while methods with overestimation bias prevail under asymmetric penalties. Our findings highlight that estimator choice should be matched to deployment context. More broadly, our work establishes activation space as a tractable domain for rare-event estimation in language models, circumventing the brittleness of discrete input-space search.
发表机构
- Johns Hopkins University(约翰·霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。