arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22538cs.AI

STAGE:基于策略范围上下文与确定性控制的面向智能体图执行的有状态翻译

STAGE: Stateful Translation to Agentic Graph Execution with Policy-Scoped Context and Deterministic Control

Mengxi Luo, Changjia Chen, An Cao, Zirong Huang, Wanyi Dai

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出可执行图框架\textsc{Stage},将模型判断限缩于策略范围节点、流程控制交由确定性代码,在多基准测试中提升了策略执行的任务成功率与可靠性。

中文摘要 AI 辅助

受策略管控的智能体必须在遵循授权流程的同时解读案件证据。我们提出了可执行图框架\textsc{Stage},该框架将模型的判断限制在策略范围节点内,同时将流程控制交由确定性代码处理。在每个节点处,模型会接收与任务相关的策略上下文并返回类型化结果,而协调器则会执行经审核的执行契约。我们在SOP-Bench Referral Abuse、两个$\tau^2$-bench领域以及专有银行基准Smart Dispute上对\textsc{Stage}进行评估。与整体式全策略执行相比,\textsc{Stage}在不同流程复杂度的工作流中普遍提升了任务成功率与重复运行可靠性。在更深层次的电信和Smart Dispute工作流中,提升最为显著,其中$\text{Pass}^3$分别提升了7.5至55.0个百分点和57.2至65.7个百分点,具体数值取决于所使用的模型。这些结果表明,将策略范围上下文与确定性流程控制相结合可提升策略执行的可靠性。

英文摘要

Policy-governed agents must interpret case evidence while reliably following authorized procedures. We present STAGE, an executable-graph framework that confines model judgment to policy-scoped nodes while placing procedural control in deterministic code. We evaluate STAGE on three public policy-following benchmarks and Smart Dispute, a proprietary banking benchmark. Compared with monolithic full-policy execution, STAGE improves task success and repeated-run reliability, with its largest observed gains on the deeper workflows. On $τ^2$-bench Telecom and Smart Dispute, $\mathrm{Pass}^{3}$ improves by up to 55.0 and 65.7 percentage points, respectively. These results demonstrate the value of combining localized policy reasoning with deterministic procedural control for enterprise use.

↑