发表机构
University of Birmingham(伯明翰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型代理作事务编译器时JSON模式等结构化输出的语义可靠性问题,引入OrderBench基准,通过对四个开放模型测试发现模式有效输出仍有语义错误率,强调结构化输出是必要接口层,不能替代域验证和故障封闭执行。
AI 中文摘要
大语言模型代理越来越多地被用作事务编译器:用户用自然语言陈述意图,模型生成一个API可执行的结构化对象。JSON模式和提供程序级别的结构化输出模式很有用,因为它们消除了大量解析失败,但它们本身并不能决定对象是否是安全、可靠的事务。我们引入了OrderBench,这是一个用于餐厅订购代理的确定性基准,它区分语法有效性、模式有效性、状态决策、精确项目语义、约束保留和不安全接受。通过对四个开放模型进行2400次Nebius Token Factory调用,采用仅提示和JSON模式模式,我们发现模式有效的输出仍可能有较大的语义错误率。在最强的模型中,两种模式都达到了100%的模式有效性,但语义成功率仍接近80%;在较弱的模型中,模式有效的不安全接受率达到两位数。结果是一个具体的工程警告:结构化输出是必要的接口层,而不是域验证和故障封闭执行的替代品。
英文摘要
LLM agents are increasingly used as transaction compilers: a user states an intent in natural language, and the model emits a structured object that an API can execute. JSON Schema and provider-level structured-output modes are useful because they remove a large class of parse failures, but they do not by themselves decide whether the object is a safe, faithful transaction. We introduce OrderBench, a deterministic benchmark for restaurant ordering agents that separates syntactic validity, schema validity, status decisions, exact item semantics, constraint preservation, and unsafe acceptances. Across 2,400 Nebius Token Factory calls to four open models in prompt-only and JSON-schema modes, we find that schema-valid output can still have large semantic error rates. In the strongest model, both modes achieve 100% schema validity, yet semantic success remains near 80%; in weaker models, schema-valid unsafe acceptances occur in double digits. The result is a concrete engineering warning: structured output is a necessary interface layer, not a substitute for domain verification and fail-closed execution.
Comments7 pages, 6 tables