arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

道德检查:针对节奏问题的战略人工智能治理

The Moral Check: Strategic AI Governance for the Pacing Problem

Zaid Amin, Rahma Santhi Zinaida, Nazlena Mohamad Ali

arXiv 2609.22869首次发表:更新:

发表机构

INTI International University; Sunway University; Universiti Kebangsaan Malaysia(英迪国际大学; 双威大学; 马来西亚国民大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对AI指数级扩展中的节奏问题,提出战略治理框架SAGE-X,通过四大思维支柱和道德检查指数确保进步不超越道德判断。

AI 中文摘要

技术无法自我引导。战略提供这种引导,确立了目的和判断必须先于计算资本的原则。随着前沿人工智能(AI)生态系统呈指数级加速发展,节奏问题导致工程团队出现严重的认知隧道效应,将标量吞吐量置于人类判断之上。校准的节奏率对于检查无节制的扩展、保障安全以及构建社会能够对其对齐性给予合理信任的模型至关重要。传统监督通过回顾性检查清单而失效,这是一种表演性治理的病态,表现出高程序成熟度但令人担忧的科学不成熟性。我们做出了双重贡献:一项PRISMA 2020综述,综合了130项实证研究(在18个基准上经MMAT评估;总语料库N=130项实证研究,涵盖178个参考基础),以及战略人工智能治理事前框架(SAGE-X)。我们的综述揭示了两个系统性漏洞:递归保证悖论(相关的、无根据的评估者信心)和持久性缺陷(在多轮转换下护栏衰减)。基于麻省理工学院战略计算原则,SAGE-X实施了四大战略思维支柱:(1)意图优于执行(缓解速度短视);(2)果断权衡(确定性绊线消除道德风险);(3)结果优于产出(审计实证危害终点);以及(4)主动对齐(将事前门控与运行时遥测同步)。在可计算的道德检查指数(MCI)及不可绕过的绊线的治理下,SAGE-X在斯坦福WebProtégé上提供了一套可操作的企业生命周期审计工具(“道德检查审计卡”),确保指数级进步永远不会超越深思熟虑的道德判断、人类能动性和社会信任。

英文摘要

Technology cannot steer itself. Strategy provides that steering, establishing the rule that purpose and judgment must precede compute capital. As the frontier artificial intelligence (AI) ecosystem accelerates exponentially, the pacing problem induces severe cognitive tunneling in engineering teams, prioritizing scalar throughput over human judgment. A calibrated pacing rate is imperative to check unchecked scaling, guarantee safety, and build models in whose alignment society can place warranted confidence. Traditional oversight fails through retrospective checklists, a pathology of performative governance exhibiting high procedural maturity but alarming scientific immaturity. We deliver a dual contribution: a PRISMA 2020 review synthesizing 130 empirical studies (MMAT-appraised across 18 benchmarks; total corpus N = 130 empirical studies across 178 reference foundations), and the Strategic AI Governance Ex-Ante Framework (SAGE-X). Our synthesis exposes two systemic vulnerabilities: the Recursive Assurance Paradox (correlated, ungrounded evaluator confidence) and the Durability Deficit (guardrail decay under multi-turn shifts). Grounded in MIT Strategic Computing doctrines, SAGE-X operationalizes Four Strategic Mindset Pillars: (1) Intent over Execution (mitigating velocity myopia); (2) Ruthless Trade-offs (deterministic tripwires eliminating moral hazard); (3) Outcomes over Outputs (auditing empirical hazard endpoints); and (4) Proactive Alignment (synchronizing ex-ante gates with runtime telemetry). Governed by a calculable Moral Check Index (MCI) with an unbypassable tripwire, SAGE-X delivers an operational Enterprise Lifecycle Audit Instrument (the "Moral Check Audit Card") on Stanford WebProtégé, ensuring exponential progress never outpaces deliberative moral judgment, human agency, and societal trust.

Comments60 pages, 9 figures, 10 tables. Preprint submitted to Computers in Human Behavior

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑