用非推理型大语言模型从自然语言生成工作流有向无环图(DAG)
Generating Workflow DAGs from Natural Language with Non-Reasoning LLMs
- Microsoft(微软公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文针对企业联络中心自然语言路由转工作流DAG问题,用神经符号分解结合确定性编译器,提升非推理LLM生成质量,在GPT-5.3-chat上与推理模型质量统计等价。
AI中文摘要:
本文解决将企业管理员编写的自然语言路由规则转换为企业联络中心可执行工作流图的问题,每个目标是带有并行分支、优先命中回退链和各分支布尔谓词的条件动作有向无环图(DAG),编码在商业路由平台的JSON方言中。研究显示,神经符号分解能让成本更低的非推理型大语言模型(LLM)以符合生产要求的质量生成复杂工作流DAG,无需昂贵的长推理模型。核心诊断发现存在发射密度瓶颈:在含635条规则的人工合成数据基准上,模型选择正确图节点的准确率很高,但随着单次发射的相互依赖节点数量增加,属性和布尔分组的错误配置会增多。因此,研究将组合图构建从模型转移到由紧凑中间表示驱动的确定性编译器,搭配学习到的注册表选择前端,聚焦相关词汇进行生成。在四种模型上,完整系统达到约89%的LLM评判有效性、约90%的精确匹配条件准确率,以及99-100%的有效JSON,同时每条规则的提示词令牌数约为整体提示词的一半。在GPT-5.3-chat上,该方法将评判有效性提升24个百分点,达到与推理型模型开箱即用质量的统计等价,不过仍存在约8个百分点的前沿差距。本文还提出结构化生成应用的部署路径和可迁移经验。
英文摘要:
This paper addresses the problem of translating natural-language routing rules written by business administrators into executable workflow graphs for enterprise contact centers. Each target is a directed acyclic graph (DAG) of conditional actions with parallel branches, hit-first fallback chains, and per-branch Boolean predicates, encoded in the JSON dialect of a commercial routing platform. We show that neuro-symbolic decomposition enables lower-cost, non-reasoning large language models to generate complex workflow DAGs at production-relevant quality without expensive extended-reasoning models. Our central diagnostic is an emission-density bottleneck: on a 635-rule benchmark of manufactured synthetic data, models select the correct graph nodes with high accuracy but increasingly misconfigure attributes and Boolean grouping as the number of interdependent nodes emitted in one pass grows. We therefore move combinatorial graph construction from the model into a deterministic compiler driven by a compact intermediate representation, with a learned registry-selection front end that focuses generation on relevant vocabulary. Across four models, the full system reaches approximately 89% LLM-judge validity, approximately 90% exact-match condition accuracy, and 99-100% valid JSON while using roughly half the per-rule prompt tokens of a monolithic prompt. On GPT-5.3-chat, the method improves judge validity by 24 percentage points and achieves statistical equivalence to a reasoning model's out-of-the-box quality, although an approximately 8-point frontier gap remains. We also present a deployment path and transferable lessons for structured-generation applications.