arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

解析器状态已掌握:面向结构化生成的结构感知KV持久化

Parser States Already Know: Structure-Conditioned KV Persistence for Structured Generation

Linze Wu, Xinrui Chen

arXiv 2608.28276首次发表:更新:

发表机构

Hangzhou Institute for Advanced Study; University of Chinese Academy of Sciences(杭州高等研究院; 中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM结构化生成中KV压缩未利用结构信号的问题,提出PASK方法,通过解析器信号优化KV持久化,在Qwen3-4B上显著提升性能并降低资源消耗。

AI 中文摘要

结构化生成是生成JSON、SQL和函数调用的大语言模型(LLM)智能体的基础,单个字段错误就可能导致下游任务失败。约束解码已通过跟踪解析器转换来强制形式有效性,这些转换能揭示生成的令牌如何参与模式关键决策,如所需字段、参数和活动语法下的结构边界。现有的KV压缩大多未利用这类与任务相关的结构信号。我们提出PASK(Parser-Aware Structural KV Persistence,解析器感知结构KV持久化),它将解析器衍生的结构转化为特定层组的KV持久化决策。PASK通过任务错误敏感度设置最低保护阈值,结合注意力输出失真分配剩余KV容量,解决模型侧KV敏感度与任务级结构化风险之间的不匹配问题。离线校准阶段将这些信号编译为持久化策略,仅保留轻量的结构条件查找在线运行。在目标总KV预算为0.33时,PASK在Qwen3-4B上的8个BFCL非实时和实时子类别中,平均性能比最强的压缩基准高出17.39个百分点;在端到端服务中,PASK的吞吐量最高提升2.2倍,TPOT降低3.3倍,同时峰值GPU内存仅为Full KV的0.53倍。

英文摘要

Structured generation underpins large language model (LLM) agents that produce JSON, SQL, and function calls, where a single wrong field can cause the downstream action to fail. Constrained decoding already tracks parser transitions to enforce formal validity, and these transitions expose how generated tokens participate in schema-critical decisions such as required fields, arguments, and structural boundaries under the active grammar. Existing KV compression largely leaves this task-relevant structural signal unused. We introduce PASK (Parser-Aware Structural KV Persistence), which turns parser-derived structure into layer-group-specific KV persistence decisions. PASK addresses the mismatch between model-side KV sensitivity and task-level structured risk by using task-error sensitivity to set minimum protection floors and attention-output distortion to allocate residual KV capacity. An offline calibration stage compiles these signals into a persistence policy, leaving only lightweight structure-conditioned lookup online. At a targe total KV budget of 0.33, PASK outperforms the strongest compressed baseline by 17.39 percentage points on average across eight BFCL non-live and Live subcategories on Qwen3-4B. In end-to-end serving, PASK achieves up to 2.2x higher throughput and 3.3x lower TPOT, while using 0.53x the peak GPU memory of Full KV.

CommentsWork in progress

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑