自然语言工作流尚未成为软件:面向可靠智能体执行的工件驱动编译
Natural-Language Workflows Are Not Software Yet: Artifact-Driven Compilation for Reliable Agent Execution
AI总结:
研究针对自然语言工作流智能体执行不可靠问题,提出Artic编译器将其转换为工件驱动工作流,经评估可提升任务解决率28个百分点,且跨模型、重复执行一致性分别提高32、56个百分点。
AI中文摘要:
自然语言工作流为智能体提供了类软件接口:领域专家可编写可复用的流程,智能体可将其作为指令执行。但这一愿景尚未实现可靠落地。工作流描述常将数据依赖关系隐含,导致执行器需推断某一步骤应使用的先前结果;智能体还可能因上下文压力无法遵循冗长或分支指令。我们提出Artic,这是一种工件驱动的工作流编译器,它将自然语言工作流转换为工件驱动的工作流,其中每个步骤声明其读写的工件、约束门控生成的工件,且显式控制转移可路由执行。该表示形式揭示了智能体执行需承担的强制负担,使编译器能识别依赖过多状态或包含复杂控制逻辑的步骤,并通过约束优化对其进行优化。为验证LLM辅助转换的有效性,Artic将忠实性检查分解为局部义务,并采用基于场景的试运行测试编译后的工作流区域是否符合源工作流。我们在来自11个真实领域工作流的488个问题实例上对Artic进行评估;与原始文本工作流相比,它将任务解决率提高了28个百分点。我们还表明,由Artic编译的工作流在跨模型和重复执行设置中,一致性分别提高了32和56个百分点。
英文摘要:
Natural-language workflows offer a software-like interface for agents: domain experts can write reusable procedures, and agents can execute them as instructions. This promise is not yet reliable. Workflow descriptions often leave data dependencies implicit, so the executor must infer which prior results a step should use; agents can also fail to follow long or branching instructions under context pressure. We propose Artic, an artifact-driven workflow compiler that transforms a natural-language workflow into an artifact-driven workflow in which each step declares the artifacts it reads and writes, constraints gate produced artifacts, and explicit control transfers route execution. This representation exposes the enforcement burden placed on agent execution, allowing the compiler to identify steps that depend on too much state or contain difficult control logic and refine them through constrained optimization. To validate the LLM-assisted transformation, Artic decomposes faithfulness checking into local obligations and uses scenario-based dry runs to test whether compiled workflow regions conform to the source workflow. We evaluate Artic on 488 problem instances from 11 real-world domain workflows; it improves task resolve rate by 28 percentage points over the original text workflow. We also show that workflows compiled by Artic are 32 and 56 percentage points more consistent in cross-model and repeated-execution setups, respectively.