arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15596cs.CR

从神经意图到加密授权:管理智能代理工作流程

From Neural Intent to Cryptographic Authorization: Securing AI-Driven Enterprise Workflows

Jiasi Weng, Jian Weng, Minrong Chen, Ming Li, Jia-Nan Liu, Zhi Li, Yue Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

研究人工智能驱动的智能代理工作流程中的安全问题,提出神经加密服务(NCS),通过神经符号设计,在大语言模型代理和特权工具间构建安全治理平面,经实验评估,能大幅降低攻击成功率,转变代理安全验证方式。

中文摘要 AI 辅助

人工智能驱动的智能代理工作流程迅速普及,正将传统政府和企业系统转变为基于语言、使用工具且日益自主的基础设施。传统密钥管理服务能验证谁可调用加密原语,但对运行时哪些工作流程步骤被授权却一无所知。我们提出神经加密服务(NCS),它是一个基于神经符号设计的主动安全治理平面,介于大语言模型代理和特权工具之间。在NCS下,不可信的神经规划器将自然语言指令编译成结构化计划草案,但无执行权。执行由确定性符号控制器控制,该控制器处理离线签名、哈希链接的指令流。NCS验证签名、递增验证哈希链、每次仅释放一个指令有效载荷,并强制代理提议的工具参数与已验证的有效载荷之间严格绑定。不匹配或乱序的工具调用被拒绝,同时保留先前验证的状态用于事后审计。我们在AgentDojo和自定义的参数劫持基准上评估了NCS。NCS将攻击成功率降至接近零,同时在良性工作流程上保持可接受的效用。因此,NCS将代理安全从询问模型意图是否合规转变为询问提议的调度是否与加密授权步骤匹配。

英文摘要

The rapid adoption of artificial intelligence (AI)-driven workflows is transforming high-consequence government and enterprise systems into language-based, tool-using and increasingly autonomous infrastructures. While these workflows can delegate planning autonomously, security-critical execution should be strictly mediated. Conventional identity management services authenticate who may invoke a primitive, but remain agnostic to which workflow steps are authorized at runtime. An AI-driven workflow can still be hijacked by injection attacks into executing malicious actions that satisfy identity checks yet violate user intent. We propose Neural Cryptographic Services (NCS), a neuro-symbolic security enforcement plane interposed between neural planners and privileged tools. NCS decouples cognitive planning from execution authority: an untrusted neural planner drafts structured plans, while a deterministic symbolic controller gates execution using an offline-signed, hash-chained instruction stream. Specifically, NCS validates cryptographic signatures and hash chains incrementally, releasing a single instruction template at a time, and admitting a tool call only when its proposed parameters satisfy the constraints of the signed template. Out-of-order or altered tool calls fail-closed, and state transitions are logged for post-hoc auditing. NCS does not attempt to prevent neural planner compromise under injection; it guarantees that a compromised planner cannot dispatch actions outside the authorization. We evaluate NCS using AgentDojo, a custom argument-hijacking dataset, adaptive adversarial instructions, and TheAgentCompany. NCS drives attack success rates to near zero while preserving acceptable utility on benign workflows.

↑