发表机构
KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PolicyGuide将领域政策编译为工作流图,通过主动验证器引导LLM智能体合规执行,在多领域提升了合规率,且工作流可跨模型迁移,攻击成功率低、程序性合规性强。
AI 中文摘要
客服类大语言模型(LLM)智能体在代表用户执行操作时必须遵守组织政策。合规失败的产生要么源于被禁止的动作(如向不符合条件的用户授予变更权限),要么源于程序性要求的遗漏(如身份验证或确认)。运行时安全机制可对高风险动作进行干预,但仅针对单一动作的检查无法引导智能体完成多步骤流程。遵循工作流的系统支持按规定流程执行,但主要目标是完成工作流而非保障智能体行为合规。PolicyGuide则将每个领域政策编译为工作流图,并在用户回合边界处调用主动验证器。验证器基于持久化的图状态协调未完成的请求,并返回符合政策要求的路径上的步骤特定修正方案。在τ²-bench的航空、零售和电信领域,使用GPT-5.4智能体及验证器时,PolicyGuide将平均Pass⁴值从0.42提升至0.62,其中在流程结构最清晰的电信领域提升最大(从0.19升至0.61)。相同工作流可迁移至Claude Sonnet 4.6和Gemini 2.5 Pro智能体。补充评估发现,在对抗性用户场景下,PolicyGuide的攻击成功率最低,且在作者设计的工作流级验证中表现出最强的程序性合规性。
英文摘要
Customer-service LLM agents must follow organizational policy when acting on a user's behalf. Compliance failures arise from either forbidden actions, such as granting an ineligible change, or omitted procedural requirements, such as identification or confirmation. Runtime safeguards can intervene on risky actions, but action-local checks do not guide an agent through a multi-step procedure. Workflow-following systems support prescribed process execution, but primarily target workflow completion rather than safeguarding agent behavior. PolicyGuide instead compiles each domain policy into a workflow graph and invokes a proactive verifier at user-turn boundaries. From persisted graph state, the verifier reconciles open requests and returns step-specific remediation along a policy-compliant path. Across the $τ^2$-bench airline, retail, and telecom domains with a GPT-5.4 agent and verifier, PolicyGuide raises mean $\mathrm{Pass}^4$ from $0.42$ to $0.62$, with the largest gain on telecom ($0.19$ to $0.61$), the most workflow-structured domain. The same workflows transfer to Claude Sonnet 4.6 and Gemini 2.5 Pro agents. Complementary evaluations find the lowest observed attack-success rate under adversarial users and the strongest procedural compliance in an author-designed workflow-level validation.
Comments26 pages, 15 figures, including appendices