发表机构
Columbia Business School(哥伦比亚商学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对智能体随机轨迹中的稀有事件,提出一种通过扰动模型权重构建重要性采样提议的新方法,利用梯度搜索和自适应正则化,在多种模型上实现超过800倍的计算效率提升。
AI 中文摘要
随着智能体以更高的自主性被部署,其随机输出轨迹中即使极其罕见的事件也可能发生并造成灾难性后果。因此,安全部署不取决于这些事件是否可能发生,而取决于它们可能发生的频率。我们研究了估计由智能体自身行为中的随机变化引起的稀有事件概率的问题。估计此类风险需要在组合上庞大的轨迹空间中进行搜索。在此情况下,朴素蒙特卡洛方法在计算上不可行,而构建有效的重要性采样(IS)提议需要对上下文相关的条件分布链进行协调更改。我们开发了一种新的IS方法,通过扰动原始模型的权重来构建提议。该提议本身是一个可微参数化的语言模型,从而能够在权重空间中进行基于梯度的搜索。我们制定了一个目标函数,结合了事件放大的可微替代项和一种自适应正则化方案,该方案动态平衡放大与估计器稳定性。我们在约1.2亿和约26亿参数的模型上评估了我们的方法,涵盖三个事件族,涉及300多个稀有事件,其稀有程度低至10^{-9},参考概率的相对标准误差小于10%。在我们最可验证的设置中,我们观察到,对于概率低于10^{-7}的事件,我们的IS估计器相比朴素蒙特卡洛实现了超过800倍的计算加权效率提升。我们的实现可在该https URL获取。
英文摘要
As agents are deployed with increased autonomy, even extremely rare events along their stochastic output trajectories can occur and prove catastrophic. Safe deployment therefore does not depend on whether these events can occur, but on how often they might. We study the problem of estimating the probability of rare events that arise from stochastic variation in the agent's own actions. Estimating this type of risk requires searching over the combinatorially vast space of trajectories. Naive Monte Carlo is computationally prohibitive in this regime, and constructing effective importance sampling (IS) proposals requires coordinated changes to a context-dependent chain of conditional distributions. We develop a new IS method that perturbs the original model's weights to construct the proposal. The proposal is itself a differentiably parameterized language model, enabling gradient-based search over weight space. We formulate an objective that combines a differentiable surrogate for event amplification and an adaptive regularization scheme that dynamically balances amplification against estimator stability. We evaluate our approach on $\sim$120M and $\sim$2.6B models across three event families spanning 300+ rare events as rare as $10^{-9}$, with reference probabilities computed with $<10\%$ relative standard error. In our most verifiable settings, we observe that our IS estimator achieves over $800\times$ compute-weighted efficiency gains over naive Monte Carlo for events with probabilities lower than $10^{-7}$. Our implementation is available at https://github.com/namkoong-lab/iterative-unalignment.