从正则表达式中分离大语言模型对齐:对抗性变异下的零覆盖率和度量相关散度
Isolating LLM Alignment from Regex: Zero Coverage and Metric-Dependent Divergence Under Adversarial Mutation
浏览论文内容
中文总结 AI 辅助
研究在语料库被设计绕过正则表达式时大语言模型对齐情况,引入$L_5$-无正则表达式并评估,发现对齐贡献度量相关,自然语言探针中无额外覆盖率提升,对抗性变体中能检测到子串分类器错过的拒绝。
中文摘要 AI 辅助
生产中的大语言模型应用通常在模型端对齐之前堆叠一个正则表达式过滤器;先前的工作发现,在活动的正则表达式过滤器后面添加一个实时Gemini后端,没有可测量的覆盖率提升。我们探讨当语料库被设计用来绕过正则表达式时,这个上限是否仍然成立。我们引入了$L_5$-无正则表达式——与$L_4$-真实(Gemini-2.5-flash、令牌预算上限、速率限制、输出清理)相同,但禁用了九模式过滤器——并针对三个子语料库(结转、正则表达式绕过、对齐隔离)中的$N = 45$个对抗性探针进行评估,通过Gemini释义和PAIR在$N = 5$次复制中放大到约1555个探针运行对。在主要子串分类器下,H1被反驳:在所有五个OWASP大语言模型前10类中,$L_5$阻止率为0%(与$L_0$相比,$\Delta\text{pp}=0$,$p = 1.00$;威尔逊上限<5%)。PAIR变体上的二级大语言模型判断度量显示阻止率为56% - 100%($p < 0.01$),表明对齐确实对对抗性构建的探针有响应——但产生的拒绝对于子串匹配来说过于细微。子语料库差异预测不成立($p = 1.00$)。对齐的贡献是度量相关的:在自然语言有害请求探针上,它在正则表达式之外没有增加观察到的覆盖率;在对抗性构建的变体上,大语言模型判断检测到子串分类器错过的拒绝。锁定的语料库、变异工件和导出脚本已发布以供复制。
英文摘要
Production LLM applications commonly stack a regex filter in front of model-side alignment; prior work found no measurable coverage gain from adding a live Gemini backend behind an active regex filter. We ask whether that ceiling holds when the corpus is \emph{designed to bypass the regex}. We introduce $L_5$-no-regex -- identical to $L_4$-real (Gemini-2.5-flash, token-budget cap, rate limit, output scrub) but with the nine-pattern filter disabled -- and evaluate it against $N{=}45$ adversarial probes across three sub-corpora (carry-forward, regex-bypass, alignment-isolate), amplified by Gemini paraphrase and PAIR to ${\sim}1{,}555$ probe-run pairs over $N{=}5$ replications. Under the primary substring classifier, H1 is refuted: $L_5$ block rate is $0,%$ across all five OWASP LLM Top-10 categories ($Δ\text{pp}{=}0$ vs.\ $L_0$, $p{=}1.00$; Wilson upper bound ${<}5,%$). A secondary LLM-judge metric on PAIR variants shows $56$--$100,%$ block rates ($p{<}0.01$), revealing alignment does respond to adversarially-framed probes -- but produces refusals too nuanced for substring matching. The sub-corpus differential prediction is not supported ($p{=}1.00$). Alignment's contribution is \emph{metric-dependent}: on natural-language harmful-request probes, it adds zero observed coverage beyond the regex; on adversarially-framed variants, an LLM judge detects refusals the substring classifier misses. The locked corpus, mutation artifacts, and export scripts are released for replication.
发表机构
- Vertex AI(顶点人工智能)
机构由 AI 辅助整理,请以论文原文为准。