AI 中文总结
研究企业人工智能代理动态能力范围界定问题,提出三源权限架构,包括基于角色的上限等,贡献含600个企业任务提示的合成数据集,经验证可推动策略优化,还发布相关内容支持对动态范围界定机制的评估。
AI 中文摘要
企业人工智能代理通常在配置时被授予静态凭证集,拥有执行每项任务可能需要的所有工具。这种持续的过度授权扩大了攻击面。我们认为能力范围界定必须遵循动态最小授权原则,并应作为一种预防机制而非检测机制。我们概述了一个实例化该原则的三源架构:基于角色的上限、任务上下文分类器和策略派生的组合禁令,以对大语言模型代理的不一致和滥用情况进行分层主动防御。该架构支持强制部署和仅观察部署。作为评估此架构的第一步,我们贡献了一个包含600个企业任务提示的合成数据集,该数据集基于多部门公司政策,在一个基于15种权限工具的分类法上标记了最低所需权限,可直接映射到可部署凭证或可执行护栏。数据集通过两遍管道构建,以避免循环,并针对60条记录/688个决策的人工审核样本进行了验证。在数据集和策略之间迭代将上限违规从46次减少到3次,减少了93%。这表明合成提示生成与策略一起开发时可推动策略优化。数据集、环境规范和生成管道已发布以支持对动态范围界定机制的评估。
英文摘要
Enterprise AI agents are typically granted static credential sets at configuration time, holding every tool the role might need for every task they perform. This persistent over-privilege expands the attack surface. We argue that capability scoping must follow a dynamic least-privilege principle and be treated as a prevention mechanism before a detection one. A credential that does not exist in an agent's context cannot be misused regardless of the agent's reasoning or evasion sophistication. We outline a three-source architecture instantiating this principle: role-based ceilings, a task-context classifier, and policy-derived combination prohibitions creating a layered proactive defense against LLM agent misalignment and misuse cases. The architecture supports both enforcing and observe-only deployment; the latter records agent permission requests inconsistent with task context, producing a behavioral signal usable in misalignment research. As a first step toward evaluating this architecture, we contribute a synthetic dataset of 600 enterprise task prompts grounded in a multi-department company policy, labeled with minimum required permissions across a 15-permission tool-based taxonomy that maps directly to deployable credentials or enforceable guardrails. The dataset is constructed via a two-pass pipeline that separates prompt generation from permission labeling to avoid circularity, and is validated against a 60-record/688 decisions human-reviewed sample (Cohen's $κ= 0.917$ pre-review and $κ= 0.967$ post-review). Iterating between dataset and policy reduced ceiling violations from 46 to 3, a 93% reduction. This shows that synthetic prompt generation can drive policy refinement when the two are developed together. The dataset, environment specification, and generation pipeline are released to support evaluation of dynamic scoping mechanisms.
CommentsPublished at the Second Workshop on Agents in the Wild: Safety, Security, and Beyond (AIWILD) at ICML 2026