发表机构
University of Amsterdam; Utrecht University(阿姆斯特丹大学; 乌得勒支大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出神经符号安全护栏架构PL-Guard,通过分离神经落地与概率符号推理,在XSTest基准上大幅降低大语言模型的不安全依从率,虽过度拒绝率略高,但提升了推理可审计性。
AI 中文摘要
大语言模型安全护栏可视为策略一致性问题:系统必须确定提示-响应对中哪些与策略相关的事实成立,以及这些事实在给定策略下的含义。常见方法包括策略提示和大语言模型作为评判者的流程,常将语义落地与策略推理任务重叠:模型既解释提示-响应对,又推理是否违反策略,这会导致对有害提示的不安全依从,或对良性请求的拒绝。为分离落地与推理角色,我们提出神经符号安全护栏架构PL-Guard:采用由谓词和ProbLog规则组成的符号策略接口,本地大语言模型使用归一化的真/假 token 分数将提示-响应对落地为谓词概率,ProbLog则对符号策略执行显式概率规则推理。在XSTest基准上,基于Qwen的离线评估器发现,带有手工整理策略的PL-Guard将基础模型的不安全依从率从22.0%降至0.5%,低于大语言模型作为评判者基线的6.0%;但该方法的过度拒绝率高于大语言模型作为评判者基线,为14.4%对5.2%。这些结果表明,将神经落地与概率符号推理分离,可揭示安全与有用性的权衡,同时使安全护栏的中间推理步骤明确且可审计。
英文摘要
Large language model guardrails can be viewed as policy-consistency problems: a system must determine which policy-relevant facts hold in a prompt-response pair and what those facts imply under a given policy. Common approaches, including policy prompting and LLM-as-a-judge pipelines, often overlap the tasks of semantic grounding and policy reasoning: the model both interprets the prompt-response pair and reasons about whether a policy has been violated. This can lead to unsafe compliance with harmful prompts, or refusals to assist benign ones. To separate grounding and reasoning roles, we propose PL-Guard, a neurosymbolic guardrail architecture. Using a symbolic policy interface consisting of predicates and ProbLog rules, a local LLM grounds prompt-response pairs into predicate probabilities using renormalized True/False token scores, while ProbLog performs explicit probabilistic rule inference over the symbolic policy. On the XSTest benchmark, an offline Qwen-based evaluator finds that PL-Guard with a hand-curated policy reduces unsafe compliance from 22.0% for the base model to 0.5%, and below the 6.0% rate of an LLM-as-a-judge baseline. This comes at the cost of higher over-refusal than the LLM-as-a-judge baseline, 14.4% versus 5.2%. These results suggest that separating neural grounding from probabilistic symbolic reasoning can expose the safety-helpfulness tradeoff while making the guardrail's intermediate reasoning steps explicit and auditable.
CommentsPreliminary version of this paper was presented at the IJCAI 2026 Workshop on Logical and Symbolic Reasoning of Large Language Models