AI 中文总结
该研究测试LLM对结构化输出模式描述的权重,发现模式影响具模型依赖性,模式设计是更强杠杆,建议将提示与模式视为统一指令表面并实证验证其放置与字段设计。
AI 中文摘要
大型语言模型(LLM)填充预定义JSON模式的结构化输出已成为数据标注和信息提取的默认机制,但模式描述也引入了第二条指令通道。我们使用跨两家厂商的10种模型配置的单字段分类任务(含临时标签),测试分类标签定义更适合放在系统提示、用户提示还是模式描述中。模式描述的表现并未始终优于基于提示的放置方式;对于无推理能力的GPT-4.1和GPT-5.4,模式放置的准确率比系统提示低11-13个百分点。不过模式并非惰性元数据:当提示与模式冲突时,错误的模式指令会导致准确率下降5-45个百分点,其中Claude Haiku 4.5的准确率从52.5%降至7%,表明模式指令可覆盖提示指令,GPT-5.5的准确率从100%降至73%。进一步,在标签字段前添加必填的中间推理字段,在存在提升空间时可使仅模式的准确率提高15-24个百分点,在所有测试案例中均超过仅系统提示的表现;该效果甚至在中等推理能力的Claude Sonnet 4.6上也成立,仅扩展思考无法产生可比增益。这表明模式设计会影响模型对字段描述中编码信息的利用效率。总体而言,这些结果说明模式的影响具有模型依赖性。实践中,系统提示仍是标签定义的安全默认选择,但更重要的原则是维护单一真实来源,防止提示/模式漂移。更关键的是,模式设计本身可能是比指令放置更强的杠杆。从业者应将提示和模式视为统一的指令表面,并针对目标模型实证验证放置方式和字段设计。
英文摘要
Structured output, where an LLM populates a predefined JSON schema, has become a default mechanism for data labeling and information extraction, but it also introduces a second instruction channel through schema descriptions. We tested whether classification-label definitions are better placed in the system prompt, user prompt, or schema description using a single-field classification task with nonce labels across ten model configurations from two vendors. Schema descriptions did not consistently outperform prompt-based placement; for GPT-4.1 and GPT-5.4 without reasoning, schema placement underperformed system prompts by 11-13 percentage points. Yet schemas are not inert metadata: when prompts and schemas conflicted, incorrect schema instructions caused accuracy drops of 5-45 points, with Claude Haiku 4.5 falling from 52.5% to 7%, indicating that schema instructions can override prompt instructions, and GPT-5.5 falling from 100% to 73%. Further, adding a required intermediate reasoning field before the label field improved schema-only accuracy by 15-24 points when headroom existed, exceeding system-prompt-only performance in every case tested. The effect held even for Claude Sonnet 4.6 at medium reasoning, where extended thinking alone did not produce a comparable gain. This suggests that schema design can affect how effectively models use information encoded in field descriptions. Overall, these results indicate that schema influence is model-dependent. In practice, the system prompt remains a safe default for definitions, but the bigger discipline is maintaining a single source of truth and preventing prompt/schema drift. More importantly, schema design itself may be a stronger lever than instruction placement. Practitioners should treat prompts and schemas as a unified instruction surface and empirically validate both placement and field design for their target model.
Comments11 pages, 7 tables