arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19816cs.CY

将理解作为前沿AI安全决策的明确且可评估的组成部分

Understanding as an Explicit and Assessable Component of Frontier AI Safety Decisions

Stephen Barrett, Robin Bloomfield, Alexandra Chirilă, Mamoon Masud, David Meredith Hardy, Phillip Mulvana

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出一种基于Assurance 2.0框架的临时方法,用于明确且可评估前沿AI安全决策中的理解,经两种场景测试,该方法可应用且能驱动工程工作。

中文摘要 AI 辅助

决策者需要充分的理解才能对复杂AI系统做出良好决策。然而,AI部署决策越来越多地在时间压力下做出,加上AI生成制品的使用,可能意味着安全案例和系统卡片的存在不再能证明存在充分的理解。我们提出的用于使理解明确且可评估的临时方法,需要生成对4个理解对象(决策、决策框架、安全论证、情境中的系统)的明确描述,以及对该理解充分性的论证。此外,该方法提供了一种机制,用于描述和评估决策者对该理解的表征的充分性。它基于安全案例的最新发展,使用Assurance 2.0框架来实施Elgin和Arendt关于理解的哲学基础。为了评估该方法,我们测试了两种不同场景:一种是通过基于角色的分析研究的机器人公司中部署AI编码智能体的策划风险;另一种是不确定性更高、对决策更关键的“若有人构建它,所有人都会死亡”(Yudkowsky和Soares提出)的论证。该测试对这两种场景的核心发现是,该方法可以应用且具有生成性:我们发现,证明理解充分性的分析(内部一致性、锚定、恰当的虚假陈述、外部一致性)驱动了工程工作。

英文摘要

Decision makers need sufficient understanding to make good decisions about training or deploying frontier AI systems. However, such decisions are increasingly made under time-pressure, and this combined with the use of AI generated artefact creation, can mean that the existence of safety cases and system cards may no longer demonstrate that sufficient understanding exists. Our provisional methodology for making understanding explicit and assessable requires the production of an explicit description of 4 objects of understanding (decision, decision-frame, safety justification, system-in-context) and a justification for the adequacy of this understanding. In addition, the methodology provides a mechanism for describing and evaluating the adequacy of the decision-maker representation of this understanding. It builds on recent developments in safety cases using the Assurance 2.0 framework to operationalise the philosophical basis of understanding from Elgin and Arendt. To assess the methodology we trialled two different scenarios. One scenario, which we investigated through role-based analysis, concerned the risk of scheming in the deployment of an AI coding agent in a robotics company and the other scenario was for the higher uncertainty, more decision-critical argument of 'If Anyone Builds It, Everyone Dies' (Yudkowsky and Soares). The trial's central finding, for these two scenarios, is that the methodology could be applied and was found to be generative: we found the analyses that justify sufficiency of understanding (internal coherence, tethering, felicitous falsehoods, external coherence) drives the engineering.

补充信息

↑