arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

人工智能安全与对齐研究的认知规范

Epistemic Norms for AI Safety and Alignment Research

Keivan Navaie

arXiv 2607.24243首次发表:更新:

AI 中文总结

探讨人工智能安全与对齐研究和主流人工智能研究的差异,基于结构化综合识别当前对齐研究的差距维度,提出包含多项内容的{\sc ECAISA},以约束安全相关研究声明的记录、检查和依赖方式,目标是可审计性。

AI 中文摘要

主流人工智能研究强调能力增长,在平均情况性能高时容忍低故障率。人工智能安全与对齐研究有不同使命:在稀疏证据、对抗动态和厚尾风险下确保灾难性故障从不发生。我们认为这两个领域在两个分析上独立的轴上存在差异,主流认知实践在这两方面都不足。基于结构化综合,我们识别出当前对齐研究的五个交叉差距维度。为解决这些差距,我们提出了{\sc ECAISA},它包含八项原则、三级评分标准等一系列内容,其治理目标是可审计性而非认证。

英文摘要

Mainstream AI research emphasises capability growth and tolerates low failure rates when average-case performance is high. AI safety and alignment research has a different mission: to ensure that catastrophic failures never occur, under sparse evidence, adversarial dynamics, and fat-tailed risk. We argue that the two domains differ along two analytically independent axes---{\it capability profile}, demonstrating the absence of hazardous behaviours rather than the presence of positive capabilities, and {\it risk profile}, bounding worst-case outcomes under fat-tailed uncertainty rather than optimising average-case performance---and that mainstream epistemic practices are inadequate on both. Building on a structured synthesis grounded in a preregistered bibliometric baseline, we identify five cross-cutting gap dimensions in current alignment research, including the near-absence of institutionalised independent verification. To address these gaps, we propose {\sc ECAISA}, an Epistemic Code for AI Safety and Alignment comprising eight principles, a three-level scoring rubric, a four-level disclosure ladder that reconciles transparency with information-hazard and commercial-confidentiality constraints, a tiered applicability scheme, an information-hazard adjudication procedure, and seven anti-gaming mechanisms. {\sc ECAISA} does not certify that any AI system is safe; it constrains how safety-relevant research claims are documented, checked, and relied upon, with auditability rather than certification as its governance target.

Comments36 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑