arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ZonoGPT:迈向验证大型GPT模型的抽象域

ZonoGPT: Towards An Abstract Domain for Verifying Large GPT Models

Hai Duong, Thanh Le, ThanhVu Nguyen

arXiv 2609.34457首次发表:更新:

发表机构

George Mason University(乔治梅森大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ZonoGPT提出一种与深度无关的抽象域,通过结构化zonotope和融合变换高效验证大型GPT模型,首次扩展到GPT-2 Medium并成功验证1,339个实例。

AI 中文摘要

基于Transformer的模型被广泛用于推理、编码和多模态智能体任务。为了对鲁棒性、安全性和公平性等理想行为提供形式化保证,神经网络验证技术在部署前证明所需属性并提供可审计的保证。然而,先前的工作仍局限于小型或受限的Transformer,且在深层模型中保持精度仍然具有挑战性。在本工作中,我们引入了ZonoGPT,一个用于验证大型Transformer的抽象域,其空间复杂度与网络深度无关。ZonoGPT使用结构化zonotope和生成器缩减机制来高效地保持相关性。为保持精度,它引入了针对Attention和LayerNorm的块特定融合变换以保留特征关系,以及针对GELU的仿射变换以保留生成器关系。这些机制使ZonoGPT成为首个验证标准架构的方法,可扩展到官方HuggingFace模型,包括GPT-2 Medium(24个块,3亿+参数),并成功验证了文本和视觉任务中的1,339个实例。

英文摘要

Transformer-based models are widely used for reasoning, coding, and multimodal agentic tasks. To provide formal assurance of desirable behaviors, such as robustness, safety, and fairness, neural network verification techniques prove required properties and provide auditable guarantees before deployment. However, prior work remains limited to small or restricted Transformers, and maintaining precision across deep models remains challenging. In this work, we introduce ZonoGpt, an abstract domain for verifying large transformers that maintains a space complexity independent of network depth. ZonoGpt uses a structured zonotope and a generator reduction mechanism to efficiently preserve correlations. To maintain precision, it introduces block-specific fused transformations for Attention and LayerNorm that retain feature relations, along with an affine transform for GELU that preserves generator relations. These mechanisms enable ZonoGpt to be the first approach to verify standard architectures, scaling to official HuggingFace models up to GPT-2 Medium (24 blocks, 300M+ parameters) and successfully verifying 1,339 instances across text and vision tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑