arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19292cs.CYcs.AIcs.HC

我们未检测到的安全故障:现代人工智能系统中隐藏的安全关键挑战的视角

The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems

Gjergji Kasneci, Enkelejda Kasneci

首次发表
浏览论文内容

中文总结 AI 辅助

研究现代人工智能系统隐藏的安全关键挑战,提出五层框架诊断隐藏风险,识别多种未被充分认识的风险模式,给出设计、治理建议及研究议程,推动人工智能安全从模型评估转向社会技术可靠性。

中文摘要 AI 辅助

当前人工智能安全讨论仍过度聚焦于明显故障,如明显危害、严重滥用和假设的灾难性场景,这并不全面。在已部署系统中,许多严重故障更隐蔽,分布在组件中且经工作流程标准化后才被视为危险。现代人工智能系统的核心安全挑战不仅在于模型是否产生有害响应,还在于社会技术系统能否保持错误可见、可争议、可控制和可恢复的条件。我们提出五层框架诊断这些隐藏风险,包括认知完整性、控制完整性、时间完整性、组织完整性和生态系统完整性。我们识别出未被充分认识的风险模式,最后给出设计、治理建议及研究议程,推动人工智能安全从以模型为中心的评估转向社会技术可靠性。

英文摘要

Current AI safety discourse still focuses disproportionately on visible failures, including obvious harms, dramatic misuse, and hypothetical catastrophic scenarios. That focus is incomplete. In deployed systems, many of the most consequential failures are quieter: plausible rather than spectacular, distributed across components rather than localized in a single output, and normalized by workflows before they are recognized as hazards. We argue that a central safety challenge in modern AI systems is increasingly not only whether a model emits a harmful response, but whether the broader socio-technical system preserves the conditions under which errors remain visible, contestable, containable, and recoverable. We propose a five-layer framework for diagnosing these hidden risks: (1) epistemic integrity, concerning whether evidence and uncertainty are represented honestly enough to support calibrated reliance; (2) control integrity, concerning whether authority, permissions, and action boundaries remain robust under attack and optimization; (3) temporal integrity, concerning whether safety holds across sessions, memory updates, and deployment drift; (4) organizational integrity, concerning whether institutions retain the capacity to audit, assign responsibility, and intervene effectively; and (5) ecosystem integrity, concerning whether AI systems preserve rather than erode the information environment on which future oversight depends. Across these layers, we identify under-recognized risk patterns, including overreliance, uncertainty and legitimacy laundering in retrieval, prompt injection, reward hacking, memory poisoning, evaluation deception, fictional human oversight, synthetic evidence pollution, and model collapse. We conclude with design and governance recommendations and a research agenda for shifting AI safety from model-centric evaluation toward socio-technical reliability.

发表机构

  • Technical University of Munich(慕尼黑技术大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑