arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.12056cs.AI

结构税:结构化输出如何影响大语言模型的性能

Structure Tax: How Structured Output affects LLMs Performance

Vineet Kumar, Kanishka, Bhuvanesh Mandora

首次发表
浏览论文内容

中文总结 AI 辅助

该研究通过多维度评估发现,结构化输出对大语言模型的影响取决于模式设计,合理设计的结构化格式可达到或超过自由形式输出性能,将核心问题从是否结构化转为如何结构化。

中文摘要 AI 辅助

将大语言模型部署到生产环境中时,通常需要将输出约束为JSON或XML等结构化格式,现有研究将由此产生的准确率损失视为一种固有的“结构税”。我们通过评估一系列模型、数据集和模式,测量任务准确率、置信度校准以及隐藏状态几何,重新审视这一说法。结果表明,该“税”取决于模式设计而非结构本身:优先推理的字段排序与自由形式输出的准确率相当或更高,而优先答案的排序则会导致准确率大幅下降,尤其在较小的模型中。格式敏感性与任务自身的结构约束成反比,保留推理顺序的模式还能提升校准效果,通过中心核对准(CKA)可观察到,在Transformer中间层中,正确与错误表示的可分性更强。我们的发现表明,设计合理的结构化格式可达到或超过自由形式输出的性能,将关键问题从“是否要结构化”重新定义为“如何结构化”,以实现最优的推理保留。

英文摘要

Deploying large language models in production often requires constraining outputs to structured formats such as JSON or XML, and prior work treats the resulting accuracy loss as an inherent `structure tax'. We re-examine this claim by evaluating a battery of models, datasets and schemas, measuring task accuracy, confidence calibration, and hidden-state geometry. The tax turns out to depend on schema design rather than on structure per se: reasoning-first field ordering matches or exceeds free-form accuracy, while answer-first ordering causes steep drops, particularly in smaller models. Format sensitivity scales inversely with a task's own structural constraints, and schemas that preserve reasoning order also improve calibration with CKA showing greater separability between correct and incorrect representations in middle transformer layers. Our findings indicate that properly designed structured formats can match or exceed free-form performance, reframing the critical question from `whether to structure' to `how to structure' for optimal reasoning preservation.

发表机构

  • PayPal Artificial Intelligence(PayPal人工智能部门)
  • PayPal

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑