arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17865cs.AI

前沿模型在行动前会寻求安全证据吗?

Do Frontier Models Seek Safety Evidence Before Acting?

Omer Tafveez

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过SAFE基准评估前沿模型在部署决策中是否主动获取安全证据,发现不同模型策略差异显著,且行为受严重性和成本主导,而非概率,揭示了模型在行动前证据获取对安全性的关键影响。

中文摘要 AI 辅助

前沿模型通常根据上下文中已有的安全信息来评估其响应方式。我们研究了一个更早的决策点:模型在行动前是否选择获取与安全相关的证据。我们引入了SAFE,一个受控基准,其中模型在做出部署决策时,可选证据在检索成本、概率、严重性和呈现方式上有所不同。在GPT-5.5、o3、Claude Opus 4.8和Claude Sonnet 4.6上,我们发现了不同的证据获取策略:Opus几乎默认进行检查,o3最倾向于跳过且对阈值最敏感,而GPT-5.5和Sonnet处于中间状态。检查行为随严重性增加而显著增强,随检索成本增加而减少,而概率对行为的影响要弱得多:将问题的陈述可能性从10%提高到70%,检查行为最多变化21个百分点。尽管存在这些差异,各模型的阶段1推理理由主要由期望值推理主导。成本-义务分解进一步表明,回避行为主要由检索摩擦和对部署收益的明确威胁驱动,而非由知晓后产生的补救义务驱动。反事实干预揭示了行为与解释之间的进一步不匹配:证据框架在检查边界附近能强烈改变决策,但很大程度上未被提及,而概率虽被频繁引用,却几乎没有因果影响。这些结果表明,部署时的安全性不仅取决于模型如何响应已知风险,还取决于它们是否获取了证明行动安全所需的证据。

英文摘要

Frontier models are often evaluated on how they respond to safety information once it is already in context. We study an earlier decision point: whether models choose to acquire safety-relevant evidence before acting. We introduce SAFE, a controlled benchmark in which models make deployment decisions with optional evidence that varies in retrieval cost, probability, severity, and presentation. Across GPT-5.5, o3, Claude Opus 4.8, and Claude Sonnet 4.6, we find distinct evidence-acquisition policies: Opus inspects nearly by default, o3 is the most skip-heavy and threshold-sensitive, and GPT-5.5 and Sonnet occupy intermediate regimes. Inspection increases strongly with severity and decreases with retrieval cost, whereas probability has much weaker behavioral influence: increasing the stated likelihood of a problem from 10% to 70% changes inspection by at most 21 percentage points. Despite these differences, Stage 1 rationales are dominated by expected-value reasoning across models. A cost-obligation decomposition further shows that avoidance is driven primarily by retrieval friction and explicit threats to the deployment payoff rather than by the remediation duties created by knowing. Counterfactual interventions reveal a further mismatch between behavior and explanation: evidence framing can strongly change decisions near the inspection boundary while going largely unmentioned, whereas probability is frequently cited despite having little causal influence. These results suggest that deployment-time safety depends not only on how models respond to known risks, but also on whether they acquire the evidence needed to know that acting is safe.

发表机构

  • University of Michigan(密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑