arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

编译,然后分页:用于过程性语言模型代理的可执行标准操作程序和能力门控运行时

Compile, Then Page: Executable SOP Programs and a Capability-Gated Runtime for Procedural LLM Agents

Chenglin Yu, Li Yin, Qingxin Fan, Ying Yu, RunyangRay Zhong, Ming Li

arXiv 2607.11346首次发表:更新:

发表机构

The Hong Kong Polytechnic University; The University of Hong Kong; Zhejiang Normal University; Research Institute for Generative AI, The Hong Kong Polytechnic University(香港理工大学; 香港大学; 浙江师范大学; 香港理工大学生成式人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对企业代理遵循SOP的问题,将SOP约束编译成可执行伪代码,用程序引导堆栈机运行,通过实验表明编译文本有益,运行时指导受能力门控,给出编译后经模型级纪律检查再启用活动框架分页的实际指导。

AI 中文摘要

企业代理必须遵循长期、有条件、安全关键的标准操作程序(SOP)。我们将机器可读的SOP约束编译成可执行伪代码,并使用程序引导(PG)堆栈机运行它们,在语言模型进行语义执行时对活动框架进行分页。一项针对六个模型的三臂SOPBench研究将表示与运行时分离:编译后的文本不会显著损害性能,在官方散文表现不佳的情况下最多可提高16.0分。运行时指导由能力门控。两个强大模型独立显示出正的七域PG对比(58:19和75:31不一致对),而弱模型则受到损害。全程序游标消融(先活动框架,保留完整程序)恢复了强大模型拒绝增益的大部分;选择性可见性增加了较小的改进。配对探测和审计测量将这种差异追溯到自发状态纪律而非重建能力。在银行方面,三个主要分支从70.4上升到86.4再到92.8,拒绝正确性达到100%。实际指导:先编译;仅在模型级纪律检查后启用活动框架分页。

英文摘要

Enterprise agents must follow long-horizon, conditional, safety-critical standard operating procedures (SOPs). We compile machine-readable SOP constraints into executable pseudo-code and run them with a program-guided (PG) stack machine that pages the active frame while an LLM performs semantic execution. A three-arm SOPBench study across six models separates representation from runtime: compiled text never significantly hurts and gains up to 16.0 points where official prose underperforms. Runtime guidance is capability-gated. Two strong models independently show positive seven-domain PG contrasts (58:19 and 75:31 discordant pairs), whereas weak models are harmed. A full-program cursor ablation (active frame first, complete program retained) recovers much of the strong-model refusal gain; selective visibility adds a smaller improvement. Paired probe and audit measurements track this divide to spontaneous state discipline rather than reconstruction ability. On Bank the three primary arms rise from 70.4 to 86.4 to 92.8, with 100% refusal correctness. Practical guidance: compile first; enable active-frame paging only after a model-level discipline check.

Comments9 pages, 3 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑