arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

withhold the Completing Chunk: 面向流式大语言模型输出的确定性配对完成护栏

Withholding the Completing Chunk: Exact Release-Boundary Equivalence for Production Streaming Guardrails

Christopher M. Frost

arXiv 2608.10279首次发表:更新:

发表机构

HEOSSI (Pte.) Ltd.(HEOSSI(私人)有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出流式LLM输出的确定性配对完成护栏,通过扫描前缀扣留配对完成块,实验验证其效果,表明它是小型固定策略的精确发布边界后盾。

AI 中文摘要

流式语言模型输出存在发布时机问题:完整响应的审核在流式文本已流出后才执行,而对部分文本的重复语义分类成本高且不稳定。我们研究一种狭窄的确定性构造,其中每个已提交的危险签名是两个词汇谓词的合取。护栏在每次发布前扫描累积前缀,并扣留使两个谓词都可观测的第一个块。在四个签名族、八个块大小和32次机制试验中,流式决策与缓冲扫描器匹配,扣留了所有配对完成块;八个单谓词对照通过。在单独的512次试验策略比较中,全前缀扫描和完整缓冲检测到所有配置的配对,512字符窗口检测到96/128个,块本地扫描检测到38/128个。固定配对标记0/338个人工生成的安全响应,检测到0/394陪审团标记的不安全响应,确认其覆盖狭窄而非一般危害。校准后的官方Llama Guard 3 1B基线将310/338个安全响应分类为安全,202/394个不安全响应分类为不安全。在16384字符响应上,重复前缀扫描器的时间在测试的块大小范围内为13.261毫秒至829.640毫秒。因此,配对完成是小型固定策略的精确发布边界后盾,而非语义审核的替代品。

英文摘要

Streaming language-model output creates an enforcement boundary: a control that detects a prohibited pattern after releasing its completing chunk cannot recall it. We study a production policy in which each ordered family is the conjunction of two regular-language predicates. Incremental matching is classical. The problem is exact composition at release time across arbitrary chunk partitions, including end-of-prefix word boundaries that can change on extension. We define an ASCII-explicit policy grammar, compile each predicate to a persistent nondeterministic finite automaton (NFA), distinguish stable from provisional assertion state, apply document-order family priority, and check the decision before releasing each chunk. We show that the resulting monitor is release-boundary equivalent to an absorbing cumulative oracle for every policy in the declared grammar. Production Python and TypeScript implementations were evaluated on 101,653 partitioned cases; a public surrogate added 100,345 cases. Both campaigns produced zero oracle, cross-runtime, or intended-family mismatches. In a frozen neutral-output profile, the memoized incremental and native-regex cumulative slopes at 64-character chunks were 0.973 and 1.976. At 16,384 characters the incremental median was 30.2 ms versus 96.6 ms for native cumulative scanning at that chunk size. Native regex remained faster at 512-character chunks (12.4 versus 29.4 ms), exposing the constant-factor crossover rather than hiding it. A shared per-stream cache cap and 129-symbol alphabet bound optimization state; the campaign peaked at 364 of 4,096 without bypass. The result is policy conformance for a deterministic backstop, not evidence of semantic safety or policy completeness.

Comments9 pages, 2 figures, 4 tables. Substantially revised version

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑