arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18351cs.CRcs.AIcs.LGcs.SYeess.SY

面向可执行终端与MCP智能体的任务条件最小权限学习

Task-Conditioned Least-Privilege Learning for Executable Terminal and MCP Agents

Alexander Tu, Michael Tu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出任务条件最小权限学习框架,通过后训练优化4B参数模型的权限选择,使智能体超额权限错误率大幅降低,安全成功率显著提升,可作为工具使用智能体的额外控制层。

中文摘要 AI 辅助

使用工具的大语言模型智能体在完成任务时可能行使用户未授予或任务不需要的权限,导致超额权限错误。仅依靠传统的权限门控系统验证智能体环境是不够的。本研究探讨通过后训练是否能让40亿参数模型在可执行终端和模型上下文协议(MCP)环境中选择任务条件权限,以补充现有措施。我们提出一种框架,其中每个动作在执行前及执行后从观测到的效果中,均会在六个风险维度接受审核,审核由确定性验证器执行,验证器对任务完成度、证据、精确状态、禁止尝试及安全成功度进行评分。结合预定义的任务特定足够权限范围,我们确定轨迹的任务特定超额权限值,并在后训练中对其进行优化。我们发现,在Qwen3.5-4B上针对1500个任务使用该框架训练后,所选种子在涵盖全部500个保留任务的2896个评估回合中达到98.48%的安全成功率,而基础策略的安全成功率为64.36%,且超额权限错误事件从4.56%降至0.79%。此外,外部测试显示能力保留及提示引导的改进,一项针对400个任务的延续研究也发现泛化证据,使超额权限事件减少6.99个百分点,同时保持原有能力。我们得出结论,通过感知最小权限的后训练学习克制,可作为可执行终端和MCP环境中使用工具的智能体的额外控制层,但无法取代权限门和沙箱。

英文摘要

Tool-using large language-model agents can complete a task while exercising authority that the user did not grant or the task does not need, causing excess-authority errors. Traditional permission gating systems alone for validating agent environments are insufficient. We study whether post-training can teach a 4B-parameter model to choose task-conditioned authority in executable terminal and Model Context Protocol (MCP) environments to complement those measures. We propose a framework where each action is audited before execution and again from observed effects along six dimensions of risk. This auditing is conducted using deterministic verifiers that score completion, evidence, exact state, prohibited attempts, and safe success. In conjunction with predefined task-specific sufficient-authority envelopes, we determine task-specific excess privilege values for trajectories, which are then optimized for in post-training. We find that after training using this framework on Qwen3.5-4B over 1,500 tasks, the selected seed reaches 98.48% safe success across 2,896 evaluation episodes spanning all 500 held-out tasks, compared with 64.36% for the base policy, and reduces excess-authority error events from 4.56% to 0.79%. Furthermore, external tests show capability retention and prompt-directed improvement. A 400 task continuation study also found evidence of generalization, reducing excess-authority events by 6.99 percentage points while maintaining previous capabilities. We conclude learned restraint through least-privilege aware post-training is therefore useful as an additional control layer for tool-using agents in executable terminal and MCP environments, but it does not replace permission gates and sandboxing.

发表机构

  • Purdue University Northwest(普渡大学西北校区)
  • Columbia University(哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑