arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

前沿人工智能的嵌入式评估

Embedded Assessments for Frontier AI

Jacob Charnock, Sophie Williams, Zaheed Kara, Markus Anderljung, Alejandro Tlaie Boria, Stephen Casper, Anka Reuel, Jonas Freund

arXiv 2609.25413首次发表:更新:

发表机构

GovAI; Pour Demain; Harvard Kennedy School; Stanford University(GovAI; Pour Demain; 哈佛肯尼迪学院; 斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出前沿AI开发者应主办嵌入式评估,赋予独立评估者内部访问权限,以深入评估内部使用风险,并建议覆盖智能体监控、安全控制与模型对齐三大领域。

AI 中文摘要

对前沿人工智能的第三方评估大多在部署前通过外部接口测试模型。但前沿人工智能模型的风险取决于其开发者如何在内部使用和管理它们。近期,前沿人工智能公司的首席执行官们承诺主办嵌入式评估。这些评估将赋予独立评估者类似员工的权限,访问开发者的内部系统、员工和文档。首先,我们认为这能够对依赖于内部系统和实践的风险进行更深入、更灵活的评估,同时在更强的安全控制下提供访问权限。接着,我们考察了关于范围、信息收集、持续时间、时机、参与条款、披露和升级的七个设计问题。我们建议前沿人工智能开发者现在就开始主办嵌入式评估,至少涵盖管理内部人工智能使用风险的三个核心领域:内部智能体监控、内部智能体安全控制与权限,以及模型对齐。为使有意义的第三方审查成为可能,评估应是持续的,评估者应至少每季度发布详细报告,并应建立明确的升级机制。这些建议旨在作为起点,还需进一步措施以实现嵌入式评估的全部潜力。

英文摘要

Third-party evaluations for frontier AI have mostly tested models through external interfaces before deployment. But the risks from frontier AI models depend on how their developers use and govern them internally. Recently, CEOs of frontier AI companies have committed to hosting embedded assessments. These assessments would give independent evaluators employee-like access to a developer's internal systems, staff, and documentation. First, we argue that this can enable deeper and more flexible assessments of risks that depend on internal systems and practices, while providing access under stronger security controls. Then, we examine seven design questions about scope, information gathering, duration, timing, terms of engagement, disclosure, and escalation. We recommend that frontier AI developers begin hosting embedded assessments now, covering at least three areas central to managing risks from internal AI use: internal agent monitoring, internal agent security controls and permissions, and model alignment. To enable meaningful third-party scrutiny, assessments should be continuous, evaluators should publish detailed reports at least quarterly, and clear escalation mechanisms should be established. These recommendations are intended as a starting point, with further steps needed to realize the full potential of embedded assessments.

Comments42 pages, 7 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑