arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03838cs.AI

LatentGuard:面向大语言模型安全防护的高效可检查潜在推理方法

LatentGuard: Efficient and Inspectable Latent Reasoning for LLM Safeguards

Zhinan Liu, Jie Li, Mingyu Kang, Jiayi Ji

AI总结:

LatentGuard是将连续潜在推理引入防护模型的高效可检查框架,其8B参数版本在提升安全判定性能的同时大幅降低推理成本,还具备良好的审计效用,为可部署的LLM安全防护提供了可行路径。

AI中文摘要:

基于推理的防护模型可提升大语言模型(LLM)的安全防护能力,但为每一次交互解码显式推理依据会导致部署成本高昂。尽管潜在推理方法通过将推理过程转移至连续状态减少了token生成,但它们在安全审核领域的应用仍未得到充分探索,且缺乏部署所需的检查接口。本文提出LatentGuard,这是一种将连续潜在推理引入防护模型的高效可检查防护框架。LatentGuard采用分阶段课程学习方法,逐步将与任务对齐的文本推理依据压缩为紧凑的潜在状态,使安全判定可直接通过连续表示进行预测。为保留可检查性,一个独立的辅助解码器按需生成紧凑的审计产物,将推理依据生成排除在标准推理路径之外。实验表明,LatentGuard-8B在GuardReasoner-8B的基础上将平均加权F1值从83.95提升至84.91,同时将关键路径推理成本从268.56个生成的推理依据token降低至1.60个潜在推理token;其审计解码器的审计效用得分为85.75,为可部署的LLM安全防护提供了高效且可检查的路径。

英文摘要:

Reasoning-based guard models improve LLM safeguards, but decoding explicit rationales for every interaction makes them costly to deploy. Although latent-reasoning methods reduce token generation by moving reasoning into continuous states, they remain underexplored for safety moderation and lack an inspection interface for deployment. In this paper, we propose LatentGuard, an efficient and inspectable safeguard framework that brings continuous latent reasoning to guard models. LatentGuard uses a staged curriculum to progressively compress task-aligned textual rationales into compact latent states, enabling safety verdicts to be predicted directly from continuous representations. To preserve inspectability, an isolated auxiliary decoder generates compact audit artifacts on demand, keeping rationale generation off the standard inference path. Experiments show that LatentGuard-8B improves mean weighted F1 from 83.95 to 84.91 over GuardReasoner-8B, while reducing critical-path reasoning cost from 268.56 generated rationale tokens to 1.60 latent reasoning tokens. Its audit decoder achieves an audit utility score of 85.75, demonstrating an efficient and inspectable path toward deployable LLM safeguards.

↑