arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33671cs.CRcs.LG

COGNIT-Guard:在显式延迟与假阳性约束下,基于异构CPU-NPU置信级联的校准独立直接决策护栏

COGNIT-Guard: Calibrated Standalone Direct-Decision Guardrails with Heterogeneous CPU-NPU Confidence Cascading under Explicit Latency and False-Positive Constraints

  • Guangdong Polytechnic Normal University(广东技术师范大学)

机构由 AI 辅助整理,请以论文原文为准。

Hao Chen

AI总结:

COGNIT-Guard提出CPU快速门控与NPU直接决策模型级联的护栏,在延迟约束下实现校准决策,降低假阳性并提升OOD鲁棒性。

AI中文摘要:

基础模型安全网关何时应生成令牌,何时应直接输出校准后的决策?我们研究用于实时预摄取安全护栏的校准独立直接决策基础模型,联合解决概率校准、双重用途假阳性控制以及显式延迟SLO下的异构CPU-NPU路由问题。预摄取护栏必须在目标LLM预填充之前筛选提示,并对良性合规查询产生低误报;然而,浅层分类器对措辞变化脆弱,隐藏状态探针需要与目标LLM耦合,而生成式护栏则产生高解码延迟和双重用途假阳性。我们提出COGNIT-Guard,将验证校准的CPU快速门控器与置信门控升级至NPU驻留的322M双向直接决策模型(Laya-322M)相结合,并采用非对称假阳性惩罚。在干净的未见DUCS-Bench测试集(N=607)上,COGNIT-Guard达到98.85%的准确率(与ML相比,McNemar p = 1.19 × 10^-4),将良性假阳性率降至0.42%(1/238;与ML相比,Fisher精确检验p = 8.23 × 10^-4),并达到1.12%的ECE和0.0104的Brier分数。在华为昇腾910C NPU上,纯NPU推理平均延迟为21.77毫秒(45.90 QPS),而实时串行CPU-NPU级联(θ*_deploy=0.70)实现41.63毫秒的平均延迟(P50:39.47毫秒,99.23%准确率,0.00%假阳性率)。在SafetyBench-ZH(N=2,100)上的评估以及与双编码器直接决策基线(CLM-8B)的比较,厘清了域内增益、OOD对齐税(Laya上从60.33%降至56.81%;域内CLM-8B为55.10%)以及经验回放恢复,将OOD准确率恢复至64.10%-65.05%,并达到99.67%-99.84%的域内准确率,假阳性率为0.00%-0.42%。

英文摘要:

When must a foundation-model safety gateway generate tokens, and when should it directly output a calibrated decision? We study calibrated standalone direct-decision foundation models for real-time pre-ingestion safety guardrails, jointly addressing probability calibration, dual-use false-positive control, and heterogeneous CPU-NPU routing under explicit latency SLOs. Pre-ingestion guardrails must screen prompts prior to target-LLM prefill with low false alarms on benign compliance inquiries; however, shallow classifiers are brittle to phrasing shifts, hidden-state probes require coupling to a target LLM, and generative guards incur high decoding latency and dual-use false positives. We present COGNIT-Guard, coupling a validation-calibrated CPU fast gatekeeper with confidence-gated escalation to an NPU-resident 322M bidirectional direct-decision model (Laya-322M) under an asymmetric false-positive penalty. On the clean unseen DUCS-Bench test split ($N=607$), COGNIT-Guard achieves 98.85% accuracy (McNemar $p = 1.19 \times 10^{-4}$ vs. ML), reduces benign FPR to 0.42% ($1/238$; Fisher's exact $p = 8.23 \times 10^{-4}$ vs. ML), and attains 1.12% ECE and 0.0104 Brier score. On Huawei Ascend 910C NPUs, pure NPU inference runs in 21.77 ms mean latency (45.90 QPS), while the live serial CPU-NPU cascade ($θ^*_{\mathrm{deploy}}=0.70$) achieves 41.63 ms mean latency (P50: 39.47 ms, 99.23% accuracy, 0.00% FPR). Evaluation on SafetyBench-ZH ($N=2,100$) and comparison against a bi-encoder direct-decision baseline (CLM-8B) disentangle in-domain gains, OOD alignment tax (60.33% $\to$ 56.81% on Laya; 55.10% on domain CLM-8B), and experience replay recovery, restoring OOD accuracy to 64.10%-65.05% and reaching 99.67%-99.84% in-domain accuracy with 0.00%-0.42% FPR.

补充信息

↑