发表机构
Vast Intelligence Lab; Southwest University; Sun Yat-sen University(星智实验室; 西南大学; 中山大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究指出将AI安全仅在训练阶段植入的范式不足,提出智能体安全应是兼具预防与证据层面的运行时契约,通过四类公开证据支撑该立场并给出研究议程。
AI 中文摘要
主流范式将AI安全视为通过RLHF(人类反馈强化学习)、DPO(直接偏好优化)或宪法AI在模型训练阶段植入的属性。我们认为,对于执行代码、修改文件、发送消息和更改数据库的自主智能体而言,这种范式在结构上存在不足。智能体安全应成为由管控框架(harness)强制执行的运行时契约,该契约具有两个互补层面:预防层面通过沙箱、权限网关、输出过滤器和轨迹监控器在危险动作发生前予以阻止;证据层面要求提供良好动作实际发生的可验证证明,以测试运行、日志捕获、文件差异和引用依据等确凿证据作为任务提交的准入条件。我们以四类公开证据支撑该立场,补充JSON文件中发布了行级协议和数据:对52起已记录的AI智能体与LLM安全事件的调查、包含31起无争议核心案例加1起有争议说明案例的错误完成审计、对12个公开智能体系统及管控框架的轨迹模式审计,以及对2023-2025年NeurIPS、ICML和ICLR录用的全部28560篇论文的标题级审计,显示训练阶段与部署阶段的出版物存在8-12倍的总体失衡。两个曾需强制执行安全的领域——计算机安全与实验科学,均形成了兼具预防与证据元素的运行时契约;具智能的AI如今面临同样压力。我们形式化了智能体轨迹模式与证据链,提出基于标准监控器组合的组合式准入命题,并概述了研究议程。具智能的AI中安全的正确单元是带有可验证证据的轨迹,而非模型。
英文摘要
The dominant paradigm treats AI safety as a property to be instilled during model training via RLHF, DPO, or Constitutional AI. We argue this is structurally insufficient for autonomous agents that execute code, mutate files, send messages, and modify databases. Agent safety should be a runtime contract enforced by the harness, and the contract has two complementary faces. The preventive face blocks dangerous actions before they happen via sandboxes, permission gates, output filters, and trajectory monitors. The evidential face requires verifiable proof that good actions actually happened, gating task submission on hard evidence such as test runs, log captures, file diffs, and citation grounding. We ground the position in four lines of public evidence, with row-level protocols and data released in the supplementary JSON files: a survey of 52 documented AI-agent and LLM safety incidents, a false-completion audit with 31 non-contested core cases plus one disputed illustrative case, a trajectory-schema audit of 12 public agent systems and harnesses, and a title-level audit of all 28,560 papers accepted at NeurIPS, ICML, and ICLR 2023-2025 showing a pooled 8-12x imbalance between training-time and deployment-time publication. Two prior communities that needed to enforce safety, computer security and the experimental sciences, converged on runtime contracts with both preventive and evidential elements; agentic AI is now under the same pressure. We formalize an Agent Trajectory Schema and Evidence Chain, state a compositional gating proposition based on standard monitor composition, and outline a research agenda. The right unit of safety in agentic AI is the trajectory-with-checkable-evidence, not the model.