自回归神经序列模型的概率模型检测
Probabilistic Model Checking of Autoregressive Neural Sequence Models
浏览论文内容
中文总结 AI 辅助
本文提出用概率模型检测流程,结合DTMC、PRISM、CEGAR等技术,量化自回归神经序列模型的约束违反概率与输入总体合规比例,在CAPP、SMILES模型上验证了该方法的有效性。
中文摘要 AI 辅助
测试集准确率无法反映部署自回归神经序列模型时的两个关键问题:被测系统(SUT)在采样可达的违反约束备选方案上分配了多少概率质量,以及输入总体中满足领域要求的比例。我们通过概率模型检测解决这两个问题。该流程从SUT的逐词生成中提取离散时间马尔可夫链(DTMC),使用PRISM模型检测器验证形式化的PCTL规范,并将每个输入的判定结果聚合为输入空间上的覆盖曲线。一个可靠性定理表明,DTMC是下近似,因此每个判定结果都能给出SUT真实可达概率的认证区间。由此得到的覆盖曲线在构造上是保守的。基于反例的抽象细化(CEGAR)循环自适应地收窄区间,最大似然算法提取最可能的证伪轨迹。两个案例研究验证了该流程:在测试准确率为100%的GPT-2计算机辅助工艺规划(CAPP)模型上,该流程量化了贪心解码隐藏但采样可达的概率质量,并确定了总体满足排序要求的最小训练比例,这两项均无法通过测试准确率报告;随后我们验证了词汇量扩大50倍的SMILES分子生成器,仅需更换外部化学有效性预言机,该流程即可识别结构完整性与化学有效性之间的差距。
英文摘要
Test-set accuracy is silent on two issues that matter when deploying autoregressive neural sequence models: how much probability mass the system under test (SUT) places on constraint-violating alternatives that are reachable under sampling and what fraction of the input population satisfies a domain requirement. We answer both with probabilistic model checking. The pipeline extracts a discrete-time Markov chain (DTMC) from the SUT's token-by-token generation, verifies formal PCTL specifications with the PRISM model checker, and aggregates the per-input verdicts into a coverage curve over the input space. A soundness theorem establishes the DTMC as an under-approximation, so every verdict yields a certified interval on the SUT's true reachability probability. The coverage built from those verdicts is, therefore, conservative by construction. A counterexample-guided abstraction refinement (CEGAR) loop adaptively tightens the interval, and a maximum-likelihood algorithm extracts the most probable falsifying trace. Two case studies exercise the pipeline. On a GPT-2 computer-aided process-planning (CAPP) model with 100% test accuracy, the pipeline quantifies the probability mass greedy decoding hides, but that is reachable with sampling; and identifies the smallest training fraction at which an ordering requirement holds population-wide, neither of which test accuracy can report. We then verify the SMILES molecular generator with a 50x larger vocabulary. The only change is an external chemical-validity oracle, and the pipeline identifies the gap between structural completeness and chemical validity.
发表机构
- Simula Research Laboratory(西穆拉研究实验室)
机构由 AI 辅助整理,请以论文原文为准。