arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01847cs.SEcs.AIcs.CL

使用LLM作为验证器推理检测模型规范中的不一致性

Detecting Inconsistencies in Model Specifications with LLM-as-Verifier Reasoning

Zichen Xie, Mrigank Pawagi, Lize Shao, Yang Hu, Wenxi Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对模型规范中的不一致性问题,提出VeriSpec方法,通过保留自然语言并使用LLM作为验证器直接审计规范文本,在OpenAI模型规范上验证了五个不一致性,精确度最高且成本最低。

中文摘要 AI 辅助

模型规范定义了大型语言模型(LLM)应如何表现,指导对齐训练、推理时行为和评估。然而,这些规范本身可能包含缺陷:两条单独合理的原则在应用于同一情境时可能规定不兼容的行为,导致没有任何响应能同时满足两者。检测此类不一致性具有挑战性。将自然语言规范形式化可能会丢失细微差别,而基于行为的测试无法可靠地区分规范缺陷与模型行为差异。我们引入了VeriSpec,这是第一种通过审计规范文本来直接检测模型规范中不一致性的方法。我们的关键见解是保留规范的自然语言形式,同时使用LLM作为验证器。VeriSpec提取结构化的、上下文感知的规则,构建主题引导的图来聚类同一权威级别下行为相关的规则,并应用LLM作为验证器的推理来检测不一致性。将VeriSpec应用于OpenAI模型规范,我们提取了405条规则,并手动验证了五个不一致性,所有这些都已报告给其开发者,他们积极响应并已启动内部讨论。与五个基线相比,VeriSpec识别出最多的已验证不一致性,实现了最高的精确度(38.5%),并且每个已验证不一致性的成本最低(11.12美元)。这些结果确立了直接规范审计作为行为对齐评估的实用补充,在缺陷影响任何模型之前从源头捕获它们。代码可在以下网址获取:https://this URL。

英文摘要

Model specifications define how large language models (LLMs) should behave, guiding alignment training, inference-time behavior, and evaluation. Yet these specifications may themselves contain defects: two individually reasonable principles may prescribe incompatible behavior when applied to the same situation, leaving no response that satisfies both. Detecting such inconsistencies is challenging. Formalizing natural-language specifications risks losing subtle distinctions, while behavior-based testing cannot reliably distinguish specification defects from differences in model behavior. We introduce VeriSpec, the first approach to directly detect inconsistencies in model specifications by auditing the specification text itself. Our key insight is to preserve the specification in natural language while using an LLM as a verifier. VeriSpec extracts structured, context-aware rules, constructs a topic-guided graph to cluster behaviorally related rules at the same authority level, and applies LLM-as-verifier reasoning to detect inconsistencies. Applying VeriSpec to the OpenAI Model Spec, we extract 405 rules and manually validate five inconsistencies, all reported to its developers, who responded positively and have initiated internal discussions. Compared with five baselines, VeriSpec identifies the most validated inconsistencies, achieves the highest precision (38.5%), and incurs the lowest cost per validated inconsistency ($11.12). These results establish direct specification auditing as a practical complement to behavioral alignment evaluation, catching defects at the source before they shape any model. The code is available at https://github.com/HIPREL-Group/VeriSpec.

发表机构

  • University of Virginia(弗吉尼亚大学)
  • University of Pennsylvania(宾夕法尼亚大学)
  • The University of Texas at Austin(德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

↑