AI 中文总结
本文评估了AI智能体基于任务的权限作用域架构,通过实现安全门控(微调RoBERTa-large)匹配Claude Haiku 4.5性能,证明任务粒度访问控制可显著减少攻击面,为智能体部署提供可部署的安全机制。
AI 中文摘要
在许多企业环境中,AI智能体的配置方式与员工自有主机相同,即在部署时固定一组静态凭据,包含员工角色可能需要的所有权限。基于角色的访问控制之所以对人类主体做出这种妥协,是因为按任务划分权限范围不可行。对于AI智能体,这种妥协使得每个凭据都暴露在外,无论当前任务是否使用它们。这些权限之后可能被受损或不对齐的智能体利用。先前的工作(Noyan,2026)将此定义为任务上下文不匹配,并提出了一种三源权限架构,包括基于角色的权限上限、任务权限分类器和基于策略的禁令,共同预先消除暴露。该工作发布了一个包含600个提示的标注数据集以评估该架构。本文通过实现安全门控来端到端地呈现该评估;一个微调的RoBERTa-large编码器,在分类质量上与少样本训练的Claude Haiku 4.5相匹配(宏F1为0.881对0.886,精确率为0.897对0.842,严重性加权残余风险为0.63对1.12)。结果表明,受信任组件无需随其监督的智能体扩展,且该控制方法的可扩展监督余量很宽。我们还提出了一种攻击面消除指标,显示仅角色上限就关闭了严重性加权表面的27.9%,而添加任务分类器则关闭了84.4%。这一差距展示了任务粒度访问控制相对于角色粒度访问控制的安全优势,而AI智能体是第一种任务粒度访问控制可强制执行的主体类型,因为它们的任务以机器可读文本形式到达。该研究将基于任务的访问控制确立为一种可测量、可能可部署的机制,用于减少智能体部署中的攻击面。
英文摘要
AI agents are provisioned the same as employee-owned hosts in many enterprise settings with a static credential set fixed at deployment which includes all permissions the employee role might ever need. Role-based access control made this compromise for human principals because scoping access per task was infeasible. For AI agents, the compromise leaves every credential standing exposed whether or not the current task uses them. These permissions can later be utilised by a compromised or misaligned agent. Prior work (Noyan, 2026) defined this as the task-context mismatch, and proposed a three-source permission architecture which includes role-based permission ceilings, a task permission classifier and policy-based prohibitions, together eliminating the exposure preemptively. The work released a 600-prompt labelled dataset to evaluate it. This paper presents that evaluation end to end by implementing the security gate; a fine-tuned RoBERTa-large encoder which matched few-shot trained Claude Haiku 4.5 on classification quality (macro-F1 0.881 against 0.886, precision 0.897 against 0.842, severity-weighted residual risk 0.63 against 1.12). The results show the trusted component does not need to scale with the agent it supervises, and the scalable-oversight margin for this control method is wide. We also propose an attack-surface elimination metric which shows the role ceiling alone closes 27.9% of the severity-weighted surface and adding the task classifier closes 84.4%. The gap displays security advantages of task-granular access control over role-granular, and AI agents are the first principal type for which the task-granular access control is enforceable because their tasks arrive as machine-readable text. The research establishes task-based access control as a measured, potentially deployable mechanism for reducing attack surface in agentic deployments.