AI 中文总结
研究如何保障人工智能代理融入系统时的安全,提出GDM人工智能控制路线图,通过保守威胁建模、基于能力的缓解措施及一系列实际防御构建多层防御体系,平衡安全与开发者速度,为内部安全提供蓝图。
AI 中文摘要
人工智能代理正在迅速加速前沿人工智能公司的工作,助力人工智能研发、网络防御和推动科学发现。随着这些代理更紧密地融入系统,释放其全部潜力需要重新思考安全方式。不应假定人工智能代理总是完美对齐,而应构建多层防御。本文提出GDM人工智能控制路线图(v0.1),这是首个针对潜在未对齐人工智能的内部安全蓝图。报告包括:采用保守威胁建模方法,假设内部部署中有追求不良目标的人工智能对手,引入基于MITRE ATT&CK的TRAIT&R分类法;基于能力的缓解措施,将特定防御措施与模型能力发展相联系,概述四个检测层级和三个预防与响应层级;一系列实际防御措施,涵盖从当前模型的低成本干预到未来模型的高级保障。人工智能控制是新兴领域,实施这些缓解措施需在安全和开发者速度间权衡,路线图将随经验和领域发展而演变。
英文摘要
AI agents are rapidly accelerating work at frontier AI companies, helping with AI R&D, cyber-defence, and advancing scientific discoveries. As these agents become more tightly integrated into our systems, unlocking their full potential requires rethinking how we do security. We should not assume that AI agents are always perfectly aligned, but should instead build in multiple layers of defence. We present the GDM AI Control Roadmap (v0.1) -- a first-of-its-kind blueprint for internal security against potentially misaligned AI. This report provides: * Threat modelling: We adopt a conservative approach to threat modelling and assume a hypothetical AI adversary pursuing undesirable goals in internal deployment. We introduce TRAIT&R, a taxonomy of tactics and techniques available to such a hypothetical AI adversary, building on the established security framework MITRE ATT&CK. * Capability-based mitigation: Because controlling more capable models requires more costly interventions, we link specific defensive measures to evolving model capabilities (such as the ability to reason opaquely or execute complex cyberattacks). As models get more powerful, our defences should escalate accordingly. We outline four Detection tiers (D1-D4) and three Prevention and Response tiers (R1-R3). * A portfolio of practical defences: We suggest 15 concrete, tiered mitigations. These range from low-cost interventions for current models (e.g., chain-of-thought monitoring, asynchronous alerts) to advanced safeguards for future models (e.g., real-time access control, system-level anomaly detection, internal activations monitoring, and shutdown infrastructure). AI control is a nascent field, and implementing these mitigations requires navigating difficult trade-offs between security and developer velocity. We expect the roadmap to evolve as we gain more experience and as the field in turn evolves.