位置与内容:分解大语言模型生成结构化输出中的结构与内容失败
Where vs What: Decomposing Structural and Content Failures in LLM-Generated Structured Outputs
浏览论文内容
中文总结 AI 辅助
该研究针对LLM生成结构化输出的错误,提出SCD框架,发现结构保真度随复杂度下降更快,进而提出SA-RLVR方法,有效提升了JSON和表格领域的值位置准确率。
中文摘要 AI 辅助
JSON和表格等结构化输出是现代基于大语言模型(LLM)系统的核心,但生成失败的评估是整体进行的,混淆了两种不同的错误模式:位置错误(正确值出现在错误位置)和值错误(预期位置出现错误值)。我们提出结构-内容分解(Structure-Content Decomposition, SCD)框架,该框架可独立衡量结构保真度和内容准确性。将SCD应用于嵌套JSON和表格任务,涉及6种模型(7B至前沿模型),我们发现一个一致现象:随着复杂度增加,结构保真度比内容准确性下降得更早且更剧烈。在最高复杂度下,即使是具备推理能力的DeepSeek-V4-Flash,也有35%的召回值出现位置错误,而Qwen2.5-7B的该比例达74%。受控消融实验表明,该模式与模型依赖语义捷径而非输出结构的拓扑理解有关。基于这些发现,我们提出SA-RLVR,将SCD指标转化为可验证的奖励,用于通过GRPO进行强化学习。SA-RLVR成功优化了不同拓扑结构下的结构寻址:它将JSON值位置准确率(VPA)从26%提升至63%,同时对未见过的模式具有泛化能力;此外,它还持续提升了表格领域的VPA,证明结构感知奖励可直接增强多领域的结构定位能力。
英文摘要
Structured outputs such as JSON and tables are central to modern LLM-based systems, yet generation failures are evaluated monolithically, conflating two distinct error modes: placement errors (correct values at wrong positions) and value errors (wrong values at intended positions). We introduce Structure-Content Decomposition (SCD), a framework that independently measures structural fidelity and content accuracy. Applying SCD to nested JSON and table tasks across six models (7B to frontier), we uncover a consistent phenomenon: structural fidelity degrades earlier and more sharply than content accuracy as complexity increases. At the highest complexity, even DeepSeek-V4-Flash (with reasoning) misplaces 35% of recalled values, while Qwen2.5-7B misplaces 74%. Controlled ablations suggest that this pattern is associated with reliance on semantic shortcuts rather than topological understanding of output structure. Based on these findings, we propose SA-RLVR, converting SCD metrics into verifiable rewards for reinforcement learning via GRPO. SA-RLVR successfully optimizes structural addressing across distinct topologies: it lifts JSON Value Placement Accuracy (VPA) from 26% to 63% while generalizing to held-out schemas; moreover, it consistently drives VPA improvements in the table domain, demonstrating that structure-aware rewards can directly enhance multi-domain structural positioning.
发表机构
- Shenzhen University(深圳大学)
- Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
机构由 AI 辅助整理,请以论文原文为准。