一步一个脚印:以LLM自主性换取流程可预测性
One Step at a Time: Trading LLM Autonomy for Process Predictability
浏览论文内容
中文总结 AI 辅助
本研究提出通过MCP逐步交付流程步骤,以自主性换取可预测性,使执行路径预先可预测且可审计,在13个领域和四个执行者上验证,显著提升流程遵循度并减少无根据回答。
中文摘要 AI 辅助
自动化运营流程的组织不仅需要正确的结果,还需要预测流程将如何运行、知道实际运行的是哪个流程,并能够逐步检查。当智能体作为执行者时,这种可预测性通常会丧失:规定的程序被放入系统提示中,只有最终答案返回。我们改为通过模型上下文协议(MCP)逐步交付程序:服务器一次释放一个步骤,智能体执行该步骤,每个步骤返回结构化的步骤输出。这以自主性换取可预测性,并且由此构造性地产生两个属性,与执行者无关。执行路径在运行前就已规定,因此流程是预先可预测的,而非事后重建;完成的步骤记录形成机器可读的执行日志,下游工具可以逐步审计和优化。在13个SOP-Bench领域和四个从前沿(Kimi K2.5)到轻量级(Ministral 3 8B)的开源权重执行者上评估了15,475次试验,我们发现逐步交付使每个执行者的执行流程都可预测和可检查,并且在小执行者时还提高了准确性。在所有四个执行者中,流程遵循度显著提高(从76-95%提高到95-99%),无根据的回答(未执行SOP而生成正确输出)几乎消失,从试验的2.1-4.5%降至0.2-0.3%(所有95%置信区间排除零);在基于提示的交付下,know_your_business上31-49%的正确回答完全绕过SOP,即使对于前沿执行者也是如此。准确性是执行者能力发挥作用的地方:轻量级执行者获得了+6.5个百分点的有根据准确性,因为外部提供流程消除了其无法承担的重建负担,而有能力的执行者则以较小的原始准确性下降换取可预测、可审计的流程。
英文摘要
Organizations automating operational processes need more than a correct outcome: they need to predict how a process will run, know which one actually ran, and inspect it step by step. When an agent is the executor that predictability is normally lost: the prescribed procedure goes into the system prompt, and only a final answer comes back. We deliver the procedure step by step over the Model Context Protocol (MCP) instead: a server releases one step at a time, the agent executes it, and each step returns a structured step_output. This trades autonomy for predictability, and two properties then follow by construction, independent of the executor. The execution path is prescribed before the run, so the process is predictable in advance rather than reconstructed afterwards; and the completed step records form a machine-readable execution log that downstream tooling can audit and optimize step by step. Evaluating 15,475 trials across 13 SOP-Bench domains and four open-weight executors from frontier (Kimi K2.5) to lightweight (Ministral 3 8B), we find step-level delivery makes the executed process predictable and inspectable for every executor, and additionally raises accuracy when the executor is small. Across all four, process adherence rises significantly (76-95% to 95-99%) and ungrounded answers (correct outputs produced without executing the SOP) near-vanish, falling from 2.1-4.5% to 0.2-0.3% of trials (all 95% CIs exclude zero); under prompt-based delivery, 31-49% of correct answers on know_your_business bypass the SOP entirely, even for the frontier executor. Accuracy is where the executor's capability enters: the lightweight executor gains +6.5pp grounded accuracy because supplying the process externally removes a reconstruction burden it cannot carry, while capable ones trade a small raw-accuracy decrement for a predictable, auditable process.
发表机构
- Amazon Web Services(亚马逊云科技)
- University of the Bundeswehr Munich(慕尼黑联邦国防军大学)
机构由 AI 辅助整理,请以论文原文为准。