AI 中文总结
该研究通过双轴评估发现,约束解码虽能完全消除小型LLM的结构错误,但语义差距随模型规模变化,且模式符合性不足以保证语义正确性。
AI 中文摘要
参数范围在0.6B-4B的小型开源大语言模型(LLMs)越来越多地被部署用于结构化输出生成(JSON、函数调用、数据提取),然而,关于约束解码(CD)在此范围内如何与模型规模相互作用,我们知之甚少。我们在三种解码条件(原生、Outlines、XGrammar)下,对来自三个家族的五个模型在14个结构化输出任务上进行了基准测试。我们引入了一个双轴评估,将结构正确性(模式有效性)与语义正确性(内容准确性)分开。我们发现,CD在所有模型中消除了所有结构失败(模式有效性从78.6%-92.9%提升到100%),但内容准确性揭示了一个持续的、尺度相关的语义差距:类型强制失败完全可以通过CD挽救,而指令语义失败(例如,多步函数调用)仍然对CD具有抵抗力。模式符合性是语义正确性的必要条件,但不是充分条件;CD的覆盖范围恰好终止于模式符合性结束之处。
英文摘要
Small open-source large language models (LLMs) in the 0.6B-4B parameter range are increasingly deployed for structured output generation (JSON, function calling, data extraction), yet little is known about how constrained decoding (CD) interacts with model scale in this regime. We benchmark five models from three families across 14 structured-output tasks under three decoding conditions (native, Outlines, XGrammar). We introduce a two-axis evaluation that separates structural correctness (schema validity) from semantic correctness (content accuracy). We find that CD eliminates all structural failures across all models (schema validity: 78.6-92.9% to 100%), but content accuracy reveals a persistent semantic gap that is scale-dependent: type coercion failures are fully CD-rescuable, while instruction-semantic failures (e.g., multi-step function calling) remain CD-resistant. Schema conformance is necessary but not sufficient for semantic correctness; CD's reach ends exactly where schema conformance ends.
Comments6 pages, ACL format Code and task suite: https://github.com/CruiseDevice/small-llm-structured-benchmark (tag v1.0)